Living library · knowledge base
A living library that keeps every idea connected to its source.
doc·ray turns your documents into a connected corpus — extracted, decomposed, and annotated, then threaded together by the entities inside them. Every person, place, date, and organization is resolved to one shared identity, so the same idea links across your whole library. Source before synthesis: every result traces back to the document it came from.
Entities are the connective tissue
Scattered mentions of the same person, place, date, or organization are recognized and resolved — AI proposes, verification decides — into one canonical identity. That identity threads documents and concepts across the whole corpus, so related work finds related work instead of sitting in separate files.
How we work
Source before synthesis
Claims, summaries, and views trace back to the documents, spans, and entities that support them — never a fluent answer you can't check.
AI proposes, verification decides
Entity resolution and enrichment may be AI-curated, but they're untrusted by default. Low-confidence matches abstain and route to review.
Human direction
Durable truth, deletion, shared identities, and catalog structure pass through explicit human or policy review — not autonomous drift.
How the library works today
Ingest & preserve
Text, layout, and metadata from PDF, EPUB, HTML, Markdown, plain text, DOCX, PPTX, and SVG. Original bytes are stored content-addressed with checksums, so provenance stays intact.
Decompose & annotate
Sentences plus a 20-layer linguistic annotation graph — lemmas, parts of speech, and dependency structure over every sentence.
Resolve entities
People, places, dates, and organizations are recognized and canonicalized under mandatory verification, so mentions become shared identities.
Curate the library
Collections, editable titles, descriptions, and tags — with an AI-suggest → human-decide metadata workflow.
Query & connect
Full-text and keyword-in-context search, a typed GraphQL API, and a read-only MCP server for AI agents.
What you get today
Provenance-preserving ingestion
PDF, EPUB, HTML, Markdown, plain text, DOCX, PPTX, and SVG in; a normalized corpus out — original bytes kept content-addressed and checksummed to their source.
20-layer annotation graph
Every sentence carries a linguistic annotation layer you can read alongside the original text.
Entity resolution & curation
Mentions collapse to canonical identities under human-reviewed verification — the connective tissue of the library.
Search & concordance
Full-text search and keyword-in-context reading, sentence by sentence, across the corpus you can see.
GraphQL API & MCP for agents
A typed GraphQL surface, plus a read-only MCP server so AI agents read the corpus — RLS-scoped to what they're allowed.
Governed for organizations
Tenant-scoped RLS, RBAC/ABAC, SSO/SCIM, audit, backups, and operator workflows for serious collections.
Where it's going Becoming · not yet built
Relationship discovery
Traverse entity-to-entity connections several degrees out — "friend-of-a-friend" — so discovery follows the graph, not just a shared identity.
Inspiration from prior art
Semantic and similarity retrieval that surfaces relevant prior work by meaning, not keywords.
Document Composer
Compose standalone documents grounded in — and cited to — source evidence, written outside the library, never into it.
Agent-driven curation
AI agents that propose library reorganizations through a governed, human-or-policy-approved write path.
Built for individual researchers, authors, and builders who want a personal living library — and for organizations that need governed ingestion, access control, auditability, and continuity. AI agents are first-class here too: they read, resolve, and curate, but never bypass human or policy review.
Start with the corpus
Browse what's already ingested, or sign in to upload and curate.