A Primordial collection stores your app’s data locally and makes its text searchable by both meaning and exact terms.
Create a Primordial Client
Import Primordial and create a client. The identifier scopes collections and their stored data, so use a stable value for the same app account or workspace.
import Primordial
let primordial = PrimordialClient(
identifier: "account-42" // Optional
)
Two PrimordialClient() with different identifiers do not share collections. If your app does not need separate
accounts or workspaces, PrimordialClient() uses the stable default identifier.
Check Semantic Search Availability
A collection is a named group of records that Primordial stores and searches together.
Collection synchronization, search, related-record lookup, and reindexing all use the
.semanticSearch capability. Check its status before presenting the feature or asking the user
to download anything.
let requirement = await primordial.ai.status(
for: .semanticSearch
)
print(requirement.isInstalled)
print(requirement.isAvailable)
print(requirement.downloadBytes)
print(requirement.requiredAvailableBytes)
isInstalled tells you whether the required AI model is already on the device.
isAvailable also checks whether the current device can use it. Primordial loads an downloaded model when a collection operation
needs it.
When a download is required, use downloadBytes and requiredAvailableBytes to explain
the download and storage requirement in your own UI before continuing.
Activate Primordial
Activate the SDK before downloading the semantic-search model or working with a collection.
try await primordial.activate(
.evaluationKey("pk_eval_your_key_here")
)
You can check primordial.isActivated when your UI needs to reflect activation state. Keep the
evaluation key out of logs and user-visible error messages.
Prepare Semantic Search
makeAvailable checks storage,
downloads embedding model or anything missing, verifies it, and reports progress as an asynchronous stream.
for try await progress in primordial.ai.makeAvailable(
.semanticSearch
) {
print(progress.fractionCompleted)
print(progress.downloadedBytes)
print(progress.totalBytes)
print(progress.phase)
}
Calling makeAvailable again reuses an existing valid download.
On iOS, add .primordialBackgroundDownloads() once to your main SwiftUI scene. See Background downloads. If the user force-quits your app,
the download is canceled.
Use Different Embedding Model
Primordial by default uses Paraphrase Multilingual MiniLM L12 v2. If you want to use different model, see Advanced API: Use a custom embedding model.
Create a Collection
A collection groups records that belong to the same searchable dataset. Opening the same collection name with the same client identifier returns access to the same durable local data.
let notes = try primordial.collection("research notes")
print(notes.name) // "research-notes"
print(notes.clientIdentifier) // "account-42"
Primordial normalizes collection names, including whitespace and capitalization. Keep the intended name stable and use separate collections when records have different purposes or retention rules.
Describe Your Data with Typed Fields
Every record has an ID and a set of fields. Your app chooses the ID, and Primordial uses it to find and update the same record later. Keep the ID non-empty, unique within the collection, and stable.
let record = PrimordialRecord(
id: "9087978swiftactors", // generated by your app
fields: [
"title": .text("Swift actors"),
"body": .markdown("## Isolation\n\nActors protect isolated mutable state."),
"kind": .keyword("guide"),
"tags": .stringList(["swift", "concurrency"]),
"priority": .integer(2),
"rating": .double(4.8),
"published": .bool(true),
"createdAt": .date(.now)
]
)
| Field type | Use it for | Semantically indexed |
|---|---|---|
.text |
Natural-language content that people should be able to find by meaning. | Yes |
.markdown |
Markdown whose headings and structure should help organize searchable chunks. | Yes |
.keyword |
A category, state, or other exact string value. | No |
.stringList |
Tags and other lists of exact string values. | No |
.integer, .double |
Counts, ranks, measurements, and numeric filters. | No |
.bool |
Flags such as published, favorite, or archived. | No |
.date |
Timestamps and date-range filters. | No |
.text and .markdown fields are split into searchable chunks (see Chunking). Every other type remains
available for filtering, display, and source information. Use the same field type for a field in every record for example, always store priority as an integer.
Add Records to the Collection
Creating a collection does not import your app’s data automatically. Your app remains the source of truth
and supplies complete PrimordialRecord values. To insert or update one record, apply an upsert:
for try await _ in notes.apply([
.upsert(record)
]) { }
Use replaceAll(with:) when you already have the complete dataset, or
synchronize(from:options:) for a durable or paged source such as SwiftData. See
Data Sync for complete insertion, update, removal, and
synchronization workflows.
Hybrid Search by Default
Pass a natural-language query to search to get the eight most relevant matches:
let matches = try await notes.search("safe shared state")
By default, Primordial combines semantic and lexical ranking, searches every text field, gives each field equal weight, applies no record filter or minimum score, and returns up to eight matches in deterministic relevance order.
Customize the Search
Choose which text fields participate, give important fields more weight, filter eligible records, choose a search strategy, discard weak matches, or change the result limit:
let matches = try await notes.search(
"safe shared state",
fields: ["title", "body"],
weights: ["title": 1.5, "body": 1.0],
filter: .field(
"kind",
.equals(.keyword("guide"))
),
minimumScore: 0.42,
limit: 8,
strategy: .hybrid
)
Use .hybrid for the default combination, .semantic for meaning-only ranking,
or .lexical for exact words and technical identifiers. Lexical-only queries do not run the
embedding model.
Hybrid search uses fusion to order matches but keeps returned scores on the semantic candidate scale. Lexical matching discounts corpus-common query terms when more informative indexed terms are present. It canonicalizes locale-supported number words and digits to the same search term, and also considers validated heading and processor context while each match still returns the exact source excerpt.
Each PrimordialMatch has a typed file or collection source. It also includes the matching
excerpt, context, stored fields, source range, and relevance score.
minimumScore accepts a finite value from 0 through 1. Primordial
compares it inclusively with the selected strategy’s score before applying field weights, then sorts
the surviving matches and applies limit. Hybrid scores use the semantic candidate scale;
lexical-only scores are normalized BM25 scores. Omit it to preserve the unfiltered behavior.
for match in matches {
print(match.source)
print(match.collection)
print(match.recordID)
print(match.field)
print(match.excerpt)
print(match.context)
print(match.score)
print(match.fields)
print(match.sourceRange)
}
Search Imported Files
Search the raw chunks created while importing files. Results include the filename, exact excerpt, Markdown heading context, score, and source range.
let files = try await primordial.files(Set(fileURLs))
let matches = try await files.search(
"safe shared state",
minimumScore: 0.42,
limit: 8,
strategy: .hybrid
)
for match in matches {
print(match.source)
print(match.excerpt)
print(match.context)
print(match.score)
}
To search files and collections together, create one source group. The same group works with answers, chat, and custom tasks.
let sources = try primordial.sources(
files: files,
collections: [notes]
)
let matches = try await sources.search("safe shared state")
.file(name:) and .collection(name:recordID:field:) identify each result. Search
returns indexed chunks; answer generation may additionally expand, select, and combine retrieved text.
Removing a file excludes it from later mixed searches while collections remain available.
Filter Search Results
Filters are typed, so their values should match the field type stored in the collection. Combine filters
with .all, .any, and .not when one condition is not enough.
let filter = PrimordialFilter.all([
.field("kind", .equals(.keyword("guide"))),
.field("priority", .lessThanOrEqual(.integer(2))),
.field("tags", .contains("swift")),
.not(.field("published", .equals(.bool(false))))
])
let matches = try await notes.search(
"structured concurrency",
filter: filter
)
| Comparison | Meaning |
|---|---|
.equals, .notEquals |
Match or exclude one typed value. |
.oneOf |
Match any value in a supplied list. |
.lessThan, .lessThanOrEqual |
Keep values below a boundary. |
.greaterThan, .greaterThanOrEqual |
Keep values above a boundary. |
.contains |
Check text, keyword, or string-list content for a string. |
Unknown fields, incompatible comparisons, invalid weights, invalid minimum scores, and non-positive result limits fail explicitly instead of silently changing the query.
Find Related Records
Use related(to:) when you already have a record and want to find similar content. Primordial
builds the query from that record’s selected text fields and leaves the source record out of the results.
let related = try await notes.related(
to: "9087978swiftactors",
fields: ["title", "body"],
filter: .field("published", .equals(.bool(true))),
minimumScore: 0.42,
limit: 4
)
The source record is removed before the final limit, so its own chunks cannot displace related records.
Inspect a Collection
Collection statistics let your app report what is stored without reading private storage details.
let statistics = try await notes.statistics()
print(statistics.recordCount)
print(statistics.fieldCount)
print(statistics.chunkCount)
print(statistics.embeddingModelID)
print(statistics.embeddingDimensions)
print(statistics.chunkingVersion)
print(statistics.storageUsage.allocatedBytes)
storageUsage includes the collection's complete managed index, including sidecar files.
It reports logical and allocated bytes without exposing private paths.
Rebuild a Collection
A collection records the embedding configuration used to build its index. If that configuration or the
index format changes or you want to use different embedding model, Primordial reports an incompatibility instead of mixing unlike embeddings. Use reindex() to rebuild
the index deliberately.
for try await progress in notes.reindex() {
print("Prepared:", progress.completedRecords, "/", progress.totalRecords)
print(progress.phase, progress.fractionCompleted)
}
Prepared-record counts advance during indexing. Earlier releases may report only stage changes.
During .embedding, completedRecords increases after each record finishes
preparation. totalRecords stays fixed. Prepared records have not yet been saved as the
new index.
For a nonempty collection, preparation accounts for the first 95% of
fractionCompleted. The .committing phase saves the index, and
.ready confirms that it is saved and ready to use. Use the phase to decide when to show
completion, rather than the record count alone.
A large record can take longer to prepare. Updates may also skip counts if your app consumes them slowly. Show the latest count instead of expecting an event for every record.
Handle Collection Errors
Invalid names, records, fields, queries, missing record IDs, incompatible indexes, and storage failures
are reported as PrimordialCollectionError. Cancellation remains Swift’s
CancellationError, so normal structured-concurrency cancellation handling still applies.
do {
let matches = try await notes.search("local AI")
print(matches)
} catch is CancellationError {
print("Search cancelled")
} catch let error as PrimordialCollectionError {
print(error.localizedDescription)
}
Advanced vector workflows
Managed collections are the recommended path for application records. If you deliberately need precomputed embeddings, multiple named representations, or direct semantic vector search, see Use a custom embedding model.