Fewer tokens. Less storage. The same answers, proven on your own documents.
Analect is an enterprise knowledge platform from ETT that decomposes unstructured documents into addressable knowledge fragments, so AI systems answer from the fragments a question needs instead of re-reading whole documents. The result, measured on our own material, is 70 to 77 percent fewer tokens on direct questions and a stored knowledge layer around 91 percent smaller than the source files.
- Measured, not asserted: every claim reproducible on your documents
- Works inside your existing hardware, models and budget cycle
- Only the fragments a question needs ever leave your estate
AI got cheaper. The bills got bigger.
Token prices fell by roughly two thirds in a single year. In the same period 73 percent of enterprises exceeded their AI cost projections and the average enterprise AI budget grew from about $1.2 million to about $7 million. 2026 is the first year the industry spends more running models than training them.
The gap between falling prices and rising bills is volume. Agents fire ten to twenty model calls per task and re-send the conversation at every step. Retrieval pipelines inflate context three to five fold. Every question pays to re-read material the organization already owns.
Blended price per million tokens
down 67%
Average enterprise AI budget
up ~6x
The cause is structural
A language model has no index into a document. To answer any question it reads everything, so every question pays for every token, and the same document is paid for again on the next question. Retrieval helps by sending less of the document, but a chunk is still prose. Compression shaves the prose. Neither changes what a unit of knowledge is.
Between 80 and 90 percent of enterprise data is unstructured, and around 90 percent of it is never analyzed. The industry's standard answer, a retrieval stack of chunks, embeddings and indexes, more than doubles the size of the estate it was meant to tame.
Change the unit of knowledge
Analect reads a document once and decomposes it, one way, into knowledge fragments: small structured statements of what the document actually asserts, each tied back to the sentence it came from. Once the knowledge is addressable, you never need the whole document again. A question retrieves the dozen fragments that hold its answer, not the eight thousand tokens around them.
Two honest lines, because they frame everything else. Decomposition does not make a document smaller; at whole-document level it makes the material larger. And if your workload is to summarize entire documents, fragments cost more, not less. The value is in targeted, repeated questioning of a stable corpus, which is what most enterprise AI work actually is.
Measured, not asserted
The token bill
On our benchmark a frontier model answering from retrieved fragments matched its own full-document accuracy on 8 of 9 questions while using 77 percent fewer tokens. A clinical guideline that cost 8,000 tokens to read whole answered a specific factual question from 355 tokens of fragments.
The storage bill
The fragment tables for a representative document came to roughly 300 kilobytes against a 3.5 megabyte source, about 91 percent smaller, for the same answerable knowledge. A retrieval stack for the same document goes the other way. Because fragments are structured data, they deduplicate, compress, cache and access-control with the tools you already run.
Our claims cover targeted questioning. Summarize-everything workloads are excluded, and we say so wherever the numbers appear.
One platform, four workflows
Every answer shows its working
Every fragment links back to the sentence it came from, so an answer assembled from fragments arrives with its evidence attached: which statements were used, from which documents, at which lines. A chunk retrieved by similarity can cite a passage; a fragment cites the exact assertion. For model-risk teams, auditors and the traceability expectations arriving with the EU AI Act, that is the difference between an answer to be trusted and an answer to be checked.
Only the fragments a question needs cross to a model. Whole documents stay inside your estate, and the transform is one way.
Next to the tools you already know
Prompt compression
Deletes low-value tokens from each request. Lossy, per request, leaves nothing behind.
Graph and retrieval RAG
Structures knowledge with a language model, expensively, on top of the source estate. Sold on accuracy; silent on cost and storage.
Gateways and routers
Cut the price per token. Analect cuts the tokens per answer. The two multiply.
Parsing and AI ETL
Feeds the chunk, embed and index pipeline. Upstream of the problem.
Deterministic decomposition, sold on the two bills that reach a CFO, with the measurement shipped as part of the product.
Where the pressures stack
Sectors where the pressures stack: enormous document estates, repetitive questioning, tight budgets and binding sovereignty rules.
Government and public services
Doing more with less is the operating condition, and the data cannot leave.
Legal and professional services
Precedent and matter files where every answer must cite its source.
Pilot, pass, play
Pilot
A fixed-fee pilot on your documents. We decompose a slice of the estate, run the proof suite, and report accuracy, tokens and storage, zeroes included.
Pass
If the numbers do not justify going further, you pass, and the pilot report is yours to keep.
Play
If they do, we productionize: ingestion at scale, routing across the model catalog, governance and the storage layer, in your environment.
Prove it on your own estate.
A fixed-fee pilot on a slice of your documents returns your numbers: accuracy, tokens and storage, with the zeroes reported alongside the wins.
Frequently asked questions
What is Analect?
Analect is an enterprise knowledge platform from ETT. It decomposes unstructured documents into addressable knowledge fragments so that AI systems answer from the fragments a question needs rather than re-reading whole documents, cutting token usage and storage while keeping every answer traceable to its source.
How much does Analect reduce AI token costs?
On our measured samples, answering direct questions from fragments used 70 to 77 percent fewer tokens than reading the source document, with matched accuracy on 8 of 9 benchmark questions. Your figures come from a pilot on your own documents.
How much storage does Analect save?
The stored knowledge layer came out around 91 percent smaller than the source documents it replaced on our measured sample. Retrieval stacks typically make an estate larger; a fragment layer makes it smaller.
Is Analect a RAG tool?
No. Retrieval sends less of a document; Analect changes the unit of knowledge itself. Fragments are structured statements, not prose chunks, and they are produced by a deterministic decomposition rather than by running a language model over the corpus.
Does Analect work with our existing models?
Yes. Analect is model-agnostic. Forecast prices workloads across roughly 200 public models from 27 providers, and Proof lets you test any of them against your documents.
Where does our data go?
Documents are decomposed inside your environment. When a question is answered, only the fragments that question needs are sent to a model. The transform is one way, so fragments cannot be turned back into the original prose.
What does a pilot involve?
A fixed-fee engagement on a slice of your estate. We decompose the documents, run the accuracy and token tests, weigh the storage and hand you a report that includes the zeroes as well as the wins. If the numbers do not justify going further, the report is yours and you walk away.
Which workloads does Analect not help with?
Summarize-everything work. At whole-document level a fragment set is larger than the source, so workloads that need the entire document every time cost more on fragments, not less. Analect's value is in targeted, repeated questioning.