Knowledge base infrastructure

One source of truth
for humans and agents

CameoDB fuses an ACID key-value store, a document model, and a full-text search engine into a single Rust binary. Agents query it over MCP and get grounded answers back. No separate index to sync, no stale results, no guessing.

Searchable Persistent MCP-native Leaderless Single binary

Built for agents

An MCP endpoint agents call directly. Iterations run at query speed, so the agent is never blocked waiting on the database.

Built for operators

One static binary, embedded storage, no coordinator to babysit. Signed releases with SBOMs attached.

Scroll normally · the rail tracks where you are · groups [] · click any segment to jump
The dual-write problem

Your index is always
a little bit wrong

The classic stack writes to the database, then asynchronously updates a search index. Between those two moments, search results disagree with the data. An agent reading during that window gets an answer that was never true.

Traditional stack

  • Two clusters to scale and operate independently.
  • Eventual-consistency lag between write and searchability.
  • Distributed transactions and messy failure recovery.

CameoDB

  • One statically compiled binary with embedded storage.
  • Atomic writes keep KV and search index perfectly in step.
  • Shared-nothing design: horizontal scaling stays boring.
Hybrid storage engine

Durably stored means
already indexed

Redb for key-value durability, a flexible document model, and Tantivy for full-text search, unified into one atomic write path. There is no window in which a document exists but cannot be found.

STAGE 01

Sequence & WAL

A monotonic sequence ID is issued; the operation is serialized into the write-ahead log inside Redb.

STAGE 02

KV insert

The document body lands in the Redb data table, giving O(log N) point retrieval by ID.

STAGE 03

Tantivy indexing

Fields are parsed and handed to the in-memory index writer for full-text availability.

STAGE 04

Dual commit

Redb commits with fsync; Tantivy performs a smart commit against its memory budget. Fully recoverable.

AI & MCP

Point your agent at it.
That is the integration

CameoDB speaks MCP natively. Add one entry to your agent's config and the knowledge base becomes a first-class tool in Claude Code, Claude Desktop, Cursor or Windsurf.

mcp.json
{
  "mcpServers": {
    "cameodb": {
      "url": "http://localhost:9480/mcp"
    }
  }
}
  • Ingest your data once: documents, logs, catalogs, tickets.
  • Connect the agent with the config block on the left.
  • Ask. Every answer is anchored to a retrievable document.

Iteration at query speed

Point lookups return in roughly 0.1 ms. An agent can probe, refine and re-query many times inside a single reasoning step instead of stalling on the database.

Distributed mesh

No leader. No quorum.
No single point to lose

Consistent hashing places data, a Kademlia DHT discovers peers, and scatter-gather fans queries out across the mesh. Nodes join and leave without an election.

CameoDB consistent-hash ring Five nodes placed around a hash ring. Each node owns the key range ending at its position. Peers link directly to one another, with no coordinator at the centre. KEY RANGE N1 N2 N3 N4 N5 NO COORDINATOR
key range owned by N2 node on the hash ring direct peer link

Consistent hashing

Data placement without central coordination: add a node and only its share of keys moves.

Kademlia DHT

Peers discover each other. There is no seed list to maintain by hand.

Scatter-gather

Queries fan out in parallel and merge on return, so no node becomes the bottleneck.

Benchmarks & provenance

Production posture
not a promise

End-to-end figures from cameodb-bench, the harness in the repository: client-observed over HTTP, not engine micro-benchmarks. Every release ships with the artifacts an enterprise review asks for.

Bulk ingest
5,400 docs/s
500 documents per request, 0.19 ms each, on a four-shard node.
Batching gain
24×
Against one document per request. 9× at a batch of 50.
Search p99
8.3 ms
At concurrency 16 on an eight-core node, 5,000 queries a second.
Single write
4.4 ms
Per document, one per request. WAL, Redb and the in-memory index add.

The harness is closed-loop: each worker waits for its answer before issuing the next request, so these are service times at a fixed concurrency rather than an open-loop SLA. Batch size trades request latency for throughput: 500 per request costs a 365 ms request to buy the 0.19 ms document.

Verifiable, not just claimed

Every artifact ships cosign-signed with a published SHA-256, and both SPDX and CycloneDX SBOMs accompany each release. Check a build yourself →

Open source

An Apache-2.0 core and MIT clients; the server is FSL-1.1-Apache-2.0 and converts to Apache-2.0 two years after release. Built on Rust 2024 with an actor-based runtime for fault isolation.

Get going

Download, run, query

Six commands from an empty directory to a query answered. The MCP endpoint comes up with the server, so an agent can ask the same question without another step.

terminal
# macOS: grab the binary
curl -sSO https://dl.cameodb.com/mac/cameodb
chmod +x ./cameodb

# start the server on :9480
./cameodb

# in a second terminal, open the client
./cameodb client -i

# load a public dataset, schema inferred
cameodb@localhost ▶ data load books \
  https://dl.cameodb.com/examples/data/booksummaries.tsv

# search what you just indexed
cameodb@localhost ▶ search books title:"Harry Potter" limit 7