compose
Synopsis
Compose index definition file(s) from a sample list produced by list. Building a new index requires a profile from profile.
USAGE
| Argument | Description |
|---|---|
INPUT_FILE |
JSONL sample list produced by list |
-o, --output-dir DIR |
Output directory for index definitions |
-n, --name TEXT |
Name of created or updated index |
-pf, --profiles-file FILE |
Profiles YAML file with index configuration (required to build a new index) |
-S, --session-id TEXT |
Session tag appended to index names (default: timestamp) |
I/O
Input: JSONL sample list (from list), profiles YAML required only for new index creation (from profile)
Output: index definition files in OUTPUT_DIR/NAME/SESSION/, with NAME.yaml as the entry point
Advanced Options
| Option | Description |
|---|---|
-pr, --profile TEXT |
Profile name to use (default: default_profile from profiles file) |
-p, --partition-count INT |
Desired number of partitions per index, 0 for automatic (default: 0) |
-b, --split-size SIZE |
Max run size (e.g. 10GB, 5000MB) before splitting samples across indices |
-m, --partition-min-size SIZE |
Minimum partition file size (e.g. 500MB, 1GB) |
-P, --partition-count-limit INT |
Upper bound on auto partition count (default: 256) |
Description
Takes a JSONL sample list (produced by list) and generates index definition
files that can be passed to plan, build or apply.
Output files are written to OUTPUT_DIR/NAME/SESSION/, where SESSION defaults to the
current timestamp if --session-id is not provided. Pass the NAME.yaml file in that
directory as the input to plan, build or apply to process the index.
Building a new index - provide --profiles-file (produced by profile).
A layout file is written to OUTPUT_DIR/NAME_layout.yaml for future updates.
Updating an existing index - omit --profiles-file. The layout file at
OUTPUT_DIR/NAME_layout.yaml is detected and loaded automatically.
If --profile is not specified, the default_profile field in the profiles file is used.
Partitioning - each Bloom filter is split into N partition files. The partition count is
determined automatically by default, or set explicitly with --partition-count. Use
--partition-min-size to enforce a minimum file size per partition, or
--partition-count-limit to cap the auto-computed count.
Splitting - when the accumulated size of samples assigned to a span exceeds --split-size,
they are distributed across multiple sub-indices rather than one. This is useful to keep
individual index files manageable for large datasets.
Examples
# Build a new index (writes layout to ./db/my_index_layout.yaml)
kmhelpers compose samples.jsonl -o ./db -n my_index -pf profiles.yaml
# Build with a session tag (output goes to ./db/my_index/my_session/)
kmhelpers compose samples.jsonl -o ./db -n my_index -pf profiles.yaml -S my_session
# Use a specific profile
kmhelpers compose samples.jsonl -o ./db -n my_index -pf profiles.yaml --profile baseline
# Override partition count
kmhelpers compose samples.jsonl -o ./db -n my_index -pf profiles.yaml --partition-count 4
# Set minimum partition size
kmhelpers compose samples.jsonl -o ./db -n my_index -pf profiles.yaml --partition-min-size 500MB
# Split large spans across multiple sub-indices
kmhelpers compose samples.jsonl -o ./db -n my_index -pf profiles.yaml --split-size 10GB
# Update an existing index (auto-detects ./db/my_index_layout.yaml)
kmhelpers compose samples.jsonl -o ./db -n my_index
See Also
list- produce the JSONL sample listprofile- produce the profiles YAML fileapply- build indices from the generated definition files- Choosing Groups and Partitions - how
-paffects query time, RAM, and storage