Query me

Evidence

What has been tested, and what has not.

Plenty of tools in this category are sold on a demo and a promise. This page is the honest state of the testing: what has been demonstrated, the actual numbers from controlled testing, what has been tried and rejected, and what real-world testing still means and when it happens.

Demonstrated in controlled testing

  • Routing: an assistant reliably opens the one correct note for a question rather than loading the whole folder.
  • Bounded reading across many document formats: only a small part of a source has to be exposed to answer a question about it.
  • Safe handling of malformed and hostile input during setup and document intake.
  • Progressive loading: short orientation first, dense detail only when the task actually needs it.
  • Updating an existing installation to a newer release without disturbing personal memory files.

Not yet established

  • Full cross-assistant execution has not been validated end to end in a single test pass.
  • Stress testing with several assistants writing at the same moment is incomplete.
  • Wider real-world testing on damaged, scanned and unusual documents is still ahead.
  • Independent blind judging of answer quality has not been done.
  • Reduction figures below describe how little source text had to be exposed; they are not a measured saving on anyone’s AI bill, and are not presented as one.
  • No claim is made that this is the first system of its kind, or that it is certified for enterprise use.

Synthetic benchmarks

The numbers behind those claims.

These come from controlled test environments built to exercise the system, not from real customer installations or production AI-provider billing. Where a number describes a tradeoff, both sides of it are reported, including the run that did not ship.

Content exposed to the model
In a synthetic routing comparison, opening the specific note that answers a question exposed 4,755 characters across 12 targeted reads — about 65% less text than the 13,419 characters across 6 reads it took to answer the same question by loading the broader source material directly. Measured in a controlled test environment built for the comparison, not on a live installation.
What that costs locally
In warm-cache local testing, doing those extra targeted reads took roughly 0.6 milliseconds longer than reading one large file directly (five repeated runs, 0.518–0.616 ms; 0.524 ms for the single large read versus 1.148 ms for the targeted reads). That is the trade: about half a millisecond more local read time for roughly two-thirds less content reaching the model. This is a synthetic-environment measurement of local file reads, not a claim about real Google Drive latency, which depends on your network and Google’s own service.
A cloud-routing approach we tested and did not ship
We measured a version that re-checked Google Drive directly on every question instead of reading from a local cache: it saved roughly 1,743 tokens of context per question but added about 12.5 seconds of latency, against a live index that itself ran about 4,891 tokens and took roughly 2.08 seconds per round trip. By our own cost estimate that is roughly 24 times more latency than the token savings were worth, so it did not ship. The numbers elsewhere on this page describe the local-cache approach that did.
Automated tests before a release ships
The current release candidate passes 65 automated tests and 5 payment-integration contract tests, a file-by-file integrity check across every distributed core file (58 files, matched by hash), and a strict internal review pass that currently reports zero errors and zero warnings. A delivered package archive was independently confirmed to contain all 59 manifested files.
Updating an existing installation
A rehearsed update from one release to the next applied 10 tracked changes with zero blockers, left existing personal memory files untouched, verified a backup before writing anything, and correctly reported no changes needed when run a second time against the same, already-updated installation.
Catching contradictions in memory
An early internal test set of 7 crafted cases — duplicate notes, direct contradictions, information that had gone stale, and similar problems — produced the 5 expected findings with no false positives and no false negatives. This is a small first test set built to exercise the reviewer, not a broad accuracy claim.

What is next

Real-world testing, and what it will mean.

Everything above was measured in environments built to test the system on purpose. That is a legitimate way to catch bugs and measure tradeoffs, and it is not the same as real-world evidence. Real-world testing, in order, means:

Outside pilots on real folders

People who are not the author, using their own documents, their own naming habits, and their own mix of tidy and messy notes.

Damaged and unusual source material

Scanned pages, corrupted files, and formats outside the ones already covered in controlled testing.

Independent blind judging

Someone other than the system’s own tests deciding whether an answer was actually right, without knowing which setup produced it.

Simultaneous multi-assistant use under load

More than one assistant writing to the same folder at the same time, past the point already exercised in controlled testing.

This page will be updated as each of these moves from planned to measured, with the same standard applied throughout: a number is published here only once it has actually been run.

Compatibility

Supported combinations.

Nothing is marked Verified until it has been tested fresh against the exact version listed.

Current status by assistant and storage surface.
Assistant or surface Role it can take Status
ChatGPT with Google Drive connectedReads and writes memory; maintains the shared structurePending verification
Claude with Google Drive connectedReads and writes memory; leaves notes for the othersPending verification
Gemini with Google Drive connectedReads memory; requests changes through the shared mailboxPending verification
Google Drive as storageThe first supported place for the folder to livePending verification

See the offers this evidence supports.

Twenty founding copies, five setup slots, one payment.