Methodology
Token-efficiency methodology
Token efficiency is meaningful only when the workflow being compared is explicit. This page documents how OfficeMaker frames comparisons between schema-led middleware generation and agentic code/render/QA loops.
What is being compared
The relevant comparison is not “OfficeMaker versus Python” as programming languages or products. It is two workflow architectures for producing an Office artifact with an LLM in the loop.
Architecture A asks the model to generate or modify document-construction code, run it, inspect errors or rendered output, and iterate. Architecture B asks the model to produce schema-led structured state while document middleware handles file construction.
Token components to count
- Initial prompt and document brief
- Schema or API instructions supplied to the model
- Generated code or structured document state
- Runtime error/debug context returned to the model
- Rendered-image or screenshot analysis tokens where used
- Revision prompts and revised output
- Any source document/data included in context
Why middleware may reduce context
Middleware can remove file-format implementation, repeated rendering instructions and deterministic document operations from the model context. For Excel, deterministic filtering can also remove irrelevant rows before reasoning.
The saving is therefore architectural: fewer tokens are needed when less work is delegated to the model. It is not a constant percentage built into OfficeMaker.
How to reproduce a comparison
- Define the same document brief and acceptance criteria for both workflows.
- Record the model and version used.
- Record the input context supplied to each workflow.
- Count prompt and completion tokens across every generation, debug and visual-QA pass.
- Record the number of iterations required to meet the same acceptance criteria.
- Publish the document size, slide/page count, data volume and any excluded steps.
- Report both total tokens and the workflow steps that caused the difference.
Excel methodology
For data-heavy Excel comparisons, record the original row/column volume and the reduced result set after deterministic filtering. Compare a full-workbook-in-context approach only if that is actually the alternative workflow.
A 500,000-row workbook filtered to 2,500 relevant rows is a useful architecture example, but the exact saving depends on values, columns, serialization and downstream reasoning.
What should not be claimed
- A fixed token-saving percentage for every document or model
- That JSON guarantees higher model quality
- That middleware eliminates the need for human QA
- That code-first libraries are inherently inferior
- That a dated estimate remains valid after the compared models, libraries or prompts change
Current published evidence
Questions
- Does OfficeMaker guarantee lower token use?
- No. Token use depends on the document, source data, model and workflow. OfficeMaker’s architectural advantage is that file construction and deterministic processing can happen outside model context.
- Can someone reproduce the benchmark?
- Yes, if the prompt, model/version, token counts, source size, iteration count and acceptance criteria are published. This methodology page defines the minimum information needed for a meaningful comparison.