Chunking
Chunking splits a long text field into smaller pieces that can be searched independently. A search can then return the relevant part of a document instead of the entire field.
.text does this automatically. Primordial keeps meaningful paragraphs as separate searchable
chunks and combines only adjacent, very short paragraphs. Long paragraphs are split at sentences, then
words, and only at individual characters when necessary. Neighboring text overlaps only when Primordial
must split one oversized paragraph.
You can also check number of generated chunks using statistics().chunkCount (see Inspect a collection). One record can therefore contribute several chunks.
Use Markdown-Aware Chunking
If your content is in Markdown format, use .markdown(markdown). Primordial uses headings to
understand where each chunk belongs and keeps meaningful paragraphs, lists, quotes, and code blocks
independently searchable, combining only adjacent blocks that are very short. Each search result exposes its heading path through
PrimordialMatch.context, for example ["Guide > Installation"]. Answers use this
context to understand the chunk, while excerpts and citations contain only the original Markdown text.
let note = PrimordialRecord(
id: "note-1",
fields: [
"title": .text("Swift concurrency"),
"body": .markdown(markdown)
]
)
Customize How Text Is Chunked
If the default chunking used by .text and .markdown does not fit your content,
use .customProcessor to choose how a field is split. Primordial includes strategies for
character windows and separators:
// 100 characters with 50-character overlap
"body": .customProcessor(text, using: .characters(size: 100, stride: 50))
The characters strategy creates fixed-size chunks; a smaller stride adds overlap.
// Split wherever this , appears.
// The marker is not included in the chunks.
"body": .customProcessor(text, using: .segments(separator: ","))
The segments strategy creates a new chunk each time it finds the separator. For example,
"Safari,Firefox" becomes the chunks "Safari" and
"Firefox".
For more control, create a custom strategy. Primordial uses its identifier and
version to remember which rules created stored chunks. Use a unique reverse-DNS identifier,
such as com.example.support-articles. If you later change the rules, increase the version so
Primordial rebuilds that field's chunks the next time you sync its records.
let strategy = PrimordialChunkingStrategy.custom(
identifier: "com.example.support-articles",
version: 1
) { input in
input.characterWindows(size: 100, stride: 100).map { range in
PrimordialChunk(
sourceRange: range,
embeddingRange: input.expanding(range, before: 50, after: 50),
context: ["Product documentation"]
)
}
}
let record = PrimordialRecord(
id: "article-1",
fields: [
"body": .customProcessor(text, using: strategy)
]
)
Pass the strategy to .customProcessor when creating the field. This strategy creates
100-character chunks. Search considers 50 extra characters on each side and the supplied context, but
returned excerpts contain only the original 100-character chunk.
Use PrimordialChunkingStrategy.custom for short rules that belong near the field. For larger
rules that you want to reuse or test independently, define a PrimordialTextChunker type:
struct BrowserChunker: PrimordialTextChunker {
static let identifier = "com.example.browser"
static let version = 1
func chunks(from input: PrimordialChunkingInput) throws -> [PrimordialChunk] {
input.segments(separatedBy: ",").map {
PrimordialChunk(sourceRange: $0.sourceRange)
}
}
}
let record = PrimordialRecord(
id: "article-1",
fields: [
"body": .customProcessor(text, using: .custom(BrowserChunker()))
]
)
Ranges use Swift Character offsets. Primordial validates the chunks before changing the index.
Sync your records again after changing custom rules. A later reindex() can reuse saved chunk
boundaries without running your custom code again.