Store & Search

Search your data by meaning and exact terms.

Turn your app’s data into an on-device search index, then combine natural-language meaning with exact words and identifiers.

A Primordial collection stores your app’s data locally and makes its text searchable by both meaning and exact terms.

Create a Primordial Client

Import Primordial and create a client. The identifier scopes collections and their stored data, so use a stable value for the same app account or workspace.

import Primordial

let primordial = PrimordialClient(
    identifier: "account-42" // Optional
)

Two PrimordialClient() with different identifiers do not share collections. If your app does not need separate accounts or workspaces, PrimordialClient() uses the stable default identifier.

Check Semantic Search Availability

A collection is a named group of records that Primordial stores and searches together.

Collection synchronization, search, related-record lookup, and reindexing all use the .semanticSearch capability. Check its status before presenting the feature or asking the user to download anything.

let requirement = await primordial.ai.status(
    for: .semanticSearch
)

print(requirement.isInstalled)
print(requirement.isAvailable)
print(requirement.downloadBytes)
print(requirement.requiredAvailableBytes)

isInstalled tells you whether the required AI model is already on the device. isAvailable also checks whether the current device can use it. Primordial loads an downloaded model when a collection operation needs it.

When a download is required, use downloadBytes and requiredAvailableBytes to explain the download and storage requirement in your own UI before continuing.

Activate Primordial

Activate the SDK before downloading the semantic-search model or working with a collection.

try await primordial.activate(
    .evaluationKey("pk_eval_your_key_here")
)

You can check primordial.isActivated when your UI needs to reflect activation state. Keep the evaluation key out of logs and user-visible error messages.

Prepare Semantic Search

makeAvailable checks storage, downloads embedding model or anything missing, verifies it, and reports progress as an asynchronous stream.

for try await progress in primordial.ai.makeAvailable(
    .semanticSearch
) {
    print(progress.fractionCompleted)
    print(progress.downloadedBytes)
    print(progress.totalBytes)
    print(progress.phase)
}

Calling makeAvailable again reuses an existing valid download.

On iOS, add .primordialBackgroundDownloads() once to your main SwiftUI scene. See Background downloads. If the user force-quits your app, the download is canceled.

Use Different Embedding Model

Primordial by default uses Paraphrase Multilingual MiniLM L12 v2. If you want to use different model, see Advanced API: Use a custom embedding model.

Create a Collection

A collection groups records that belong to the same searchable dataset. Opening the same collection name with the same client identifier returns access to the same durable local data.

let notes = try primordial.collection("research notes")

print(notes.name) // "research-notes"
print(notes.clientIdentifier) // "account-42"

Primordial normalizes collection names, including whitespace and capitalization. Keep the intended name stable and use separate collections when records have different purposes or retention rules.

Describe Your Data with Typed Fields

Every record has an ID and a set of fields. Your app chooses the ID, and Primordial uses it to find and update the same record later. Keep the ID non-empty, unique within the collection, and stable.

let record = PrimordialRecord(
    id: "9087978swiftactors", // generated by your app
    fields: [
        "title": .text("Swift actors"),
        "body": .markdown("## Isolation\n\nActors protect isolated mutable state."),
        "kind": .keyword("guide"),
        "tags": .stringList(["swift", "concurrency"]),
        "priority": .integer(2),
        "rating": .double(4.8),
        "published": .bool(true),
        "createdAt": .date(.now)
    ]
)
Field type Use it for Semantically indexed
.text Natural-language content that people should be able to find by meaning. Yes
.markdown Markdown whose headings and structure should help organize searchable chunks. Yes
.keyword A category, state, or other exact string value. No
.stringList Tags and other lists of exact string values. No
.integer, .double Counts, ranks, measurements, and numeric filters. No
.bool Flags such as published, favorite, or archived. No
.date Timestamps and date-range filters. No

.text and .markdown fields are split into searchable chunks (see Chunking). Every other type remains available for filtering, display, and source information. Use the same field type for a field in every record for example, always store priority as an integer.

Add Records to the Collection

Creating a collection does not import your app’s data automatically. Your app remains the source of truth and supplies complete PrimordialRecord values. To insert or update one record, apply an upsert:

for try await _ in notes.apply([
    .upsert(record)
]) { }

Use replaceAll(with:) when you already have the complete dataset, or synchronize(from:options:) for a durable or paged source such as SwiftData. See Data Sync for complete insertion, update, removal, and synchronization workflows.

Hybrid Search by Default

Pass a natural-language query to search to get the eight most relevant matches:

let matches = try await notes.search("safe shared state")

By default, Primordial combines semantic and lexical ranking, searches every text field, gives each field equal weight, applies no record filter or minimum score, and returns up to eight matches in deterministic relevance order.

Customize the Search

Choose which text fields participate, give important fields more weight, filter eligible records, choose a search strategy, discard weak matches, or change the result limit:

let matches = try await notes.search(
    "safe shared state",
    fields: ["title", "body"],
    weights: ["title": 1.5, "body": 1.0],
    filter: .field(
        "kind",
        .equals(.keyword("guide"))
    ),
    minimumScore: 0.42,
    limit: 8,
    strategy: .hybrid
)

Use .hybrid for the default combination, .semantic for meaning-only ranking, or .lexical for exact words and technical identifiers. Lexical-only queries do not run the embedding model.

Hybrid search uses fusion to order matches but keeps returned scores on the semantic candidate scale. Lexical matching discounts corpus-common query terms when more informative indexed terms are present. It canonicalizes locale-supported number words and digits to the same search term, and also considers validated heading and processor context while each match still returns the exact source excerpt.

Each PrimordialMatch has a typed file or collection source. It also includes the matching excerpt, context, stored fields, source range, and relevance score.

minimumScore accepts a finite value from 0 through 1. Primordial compares it inclusively with the selected strategy’s score before applying field weights, then sorts the surviving matches and applies limit. Hybrid scores use the semantic candidate scale; lexical-only scores are normalized BM25 scores. Omit it to preserve the unfiltered behavior.

for match in matches {
    print(match.source)
    print(match.collection)
    print(match.recordID)
    print(match.field)
    print(match.excerpt)
    print(match.context)
    print(match.score)
    print(match.fields)
    print(match.sourceRange)
}

Search Imported Files

Search the raw chunks created while importing files. Results include the filename, exact excerpt, Markdown heading context, score, and source range.

let files = try await primordial.files(Set(fileURLs))

let matches = try await files.search(
    "safe shared state",
    minimumScore: 0.42,
    limit: 8,
    strategy: .hybrid
)

for match in matches {
    print(match.source)
    print(match.excerpt)
    print(match.context)
    print(match.score)
}

To search files and collections together, create one source group. The same group works with answers, chat, and custom tasks.

let sources = try primordial.sources(
    files: files,
    collections: [notes]
)

let matches = try await sources.search("safe shared state")

.file(name:) and .collection(name:recordID:field:) identify each result. Search returns indexed chunks; answer generation may additionally expand, select, and combine retrieved text. Removing a file excludes it from later mixed searches while collections remain available.

Filter Search Results

Filters are typed, so their values should match the field type stored in the collection. Combine filters with .all, .any, and .not when one condition is not enough.

let filter = PrimordialFilter.all([
    .field("kind", .equals(.keyword("guide"))),
    .field("priority", .lessThanOrEqual(.integer(2))),
    .field("tags", .contains("swift")),
    .not(.field("published", .equals(.bool(false))))
])

let matches = try await notes.search(
    "structured concurrency",
    filter: filter
)
Comparison Meaning
.equals, .notEquals Match or exclude one typed value.
.oneOf Match any value in a supplied list.
.lessThan, .lessThanOrEqual Keep values below a boundary.
.greaterThan, .greaterThanOrEqual Keep values above a boundary.
.contains Check text, keyword, or string-list content for a string.

Unknown fields, incompatible comparisons, invalid weights, invalid minimum scores, and non-positive result limits fail explicitly instead of silently changing the query.

Find Related Records

Use related(to:) when you already have a record and want to find similar content. Primordial builds the query from that record’s selected text fields and leaves the source record out of the results.

let related = try await notes.related(
    to: "9087978swiftactors",
    fields: ["title", "body"],
    filter: .field("published", .equals(.bool(true))),
    minimumScore: 0.42,
    limit: 4
)

The source record is removed before the final limit, so its own chunks cannot displace related records.

Inspect a Collection

Collection statistics let your app report what is stored without reading private storage details.

let statistics = try await notes.statistics()

print(statistics.recordCount)
print(statistics.fieldCount)
print(statistics.chunkCount)
print(statistics.embeddingModelID)
print(statistics.embeddingDimensions)
print(statistics.chunkingVersion)
print(statistics.storageUsage.allocatedBytes)

storageUsage includes the collection's complete managed index, including sidecar files. It reports logical and allocated bytes without exposing private paths.

Rebuild a Collection

A collection records the embedding configuration used to build its index. If that configuration or the index format changes or you want to use different embedding model, Primordial reports an incompatibility instead of mixing unlike embeddings. Use reindex() to rebuild the index deliberately.

for try await progress in notes.reindex() {
    print("Prepared:", progress.completedRecords, "/", progress.totalRecords)
    print(progress.phase, progress.fractionCompleted)
}

Prepared-record counts advance during indexing. Earlier releases may report only stage changes.


During .embedding, completedRecords increases after each record finishes preparation. totalRecords stays fixed. Prepared records have not yet been saved as the new index.

For a nonempty collection, preparation accounts for the first 95% of fractionCompleted. The .committing phase saves the index, and .ready confirms that it is saved and ready to use. Use the phase to decide when to show completion, rather than the record count alone.

A large record can take longer to prepare. Updates may also skip counts if your app consumes them slowly. Show the latest count instead of expecting an event for every record.

Handle Collection Errors

Invalid names, records, fields, queries, missing record IDs, incompatible indexes, and storage failures are reported as PrimordialCollectionError. Cancellation remains Swift’s CancellationError, so normal structured-concurrency cancellation handling still applies.

do {
    let matches = try await notes.search("local AI")
    print(matches)
} catch is CancellationError {
    print("Search cancelled")
} catch let error as PrimordialCollectionError {
    print(error.localizedDescription)
}

Advanced vector workflows

Managed collections are the recommended path for application records. If you deliberately need precomputed embeddings, multiple named representations, or direct semantic vector search, see Use a custom embedding model.