The catalog
Six components are visible in the experiment editor:Column stats
Per-column distributions, missingness, cardinality, and basic numeric statistics. Runs automatically on first ingest.
Predictors (EBM)
Fits an Explainable Boosting Machine and writes feature importance, shape functions, and pairwise interactions.
Correlation-association matrix
Correlations and other measures of association for every column pair.
Beta
Forecast
Time-series projection with 95% prediction intervals, using a Holt-Winters model under the hood.
Outlier report
Identifies unusual values and anomalous rows.
Beta
Feature metadata
LLM-generated, business-friendly descriptions for every dataset column. Powers chat grounding.
UMAP embedding
A 2-D projection of feature space for cluster visualization and “similar rows” queries.
Missing-value patterns
Identifies patterns in the missing data.
Beta
Distribution profile
Per-column information about the shape of the data distribution.
Beta
ebm_model, forecast_model) exist behind the scenes as dependencies — they hold the heavy fitted state and are consumed by the user-visible ones. You don’t pick them in the editor; they auto-run when their dependent component is selected.
Anatomy of a component
Every component declares the same metadata:
When you select a component in the experiment editor, the catalog is what tells the form which inputs to render and how to validate them.
How components run
Reading component outputs
Three places where outputs surface:- Dataset detail → Components tab. Per-component status plus the latest artifact, rendered with either the component’s declared blocks or a bespoke React viewer (for EBM, UMAP, and Feature metadata).
- Summand chat. The
analyzetool resolves the latest version of any component and returns its data inline. Filtering by column or feature name is supported through each component’sagent_config. - Downstream views. Components write to S3 paths and DynamoDB items you can join against in custom SQL views.
What’s not a component
A few analyses run outside the component catalog:- Surprise finding has its own Step Functions pipeline that runs alongside the semantic-layer dispatcher. It’s not selectable in the experiment editor today — surprises surface through their own page in the product.
- Sync / curation (the read-from-source-and-write-to-Parquet step) is part of the connector pipeline, not the component catalog.
- Chat itself isn’t a component — it’s a separate Lambda (
claude-stream) that consumes component outputs via theanalyzetool.
Catalog stability
The catalog is small and slow-moving. Most recent commits are bug fixes (cache improvements, edge-case handling) rather than new components. If you have a need for analysis the catalog doesn’t cover, custom components are on the roadmap for Enterprise — email enterprise@summand.com and tell us what you’d build.Per-component details
- Column stats — what runs by default on every dataset.
- Predictors (EBM) — interpretable model fitting, with shape functions and pairwise interactions.
- Correlation-association matrix — meaasures of association between every column pair.
- Forecast — time-series projection with prediction intervals.
- Feature metadata — LLM-generated column descriptions.
- UMAP embedding — 2-D projection for clustering and similarity.