Echo: Helping Old Ideas Return at the Right Moment
Project Echo · local-first Chrome extension
Type Independent AI product engineering project
Role Product definition, interaction design, development, dogfooding, evaluation, and ship decisions
Stack Plasmo · React · TypeScript · Chrome Manifest V3 · Dexie / IndexedDB
Source GitHub repository
1. Overview
I often leave useful judgments inside long conversations with ChatGPT, Claude, Gemini, or Grok. Weeks later, a new discussion may make one of those thoughts relevant again, but I rarely remember which chat contains it. Saving more text was easy. Bringing back the right thought, at a useful moment, without interrupting me with loose matches was the harder problem.
Echo is a Chrome side panel for that workflow. I select a passage, save my own thought with its source, and continue working in the same browser tab. When I later select something related, Echo can show a previous thought and a route back to its source. When the evidence is weak, it stays quiet. Records remain in the local browser profile; the core capture and recall loop does not call a model API or application backend.
This project changed direction twice through use and evaluation. A separate app called Collect was too far from the moment I wanted to save an idea. After moving capture into the browser, an early lexical recall version surfaced too many weak matches. Those failures shaped both the product and the way I evaluated it.
| What I built | What I learned | Current decision |
|---|---|---|
| A local Chrome side panel for capture, search, and contextual recall | The cost of an irrelevant suggestion matters as much as a missed memory | Keep the conservative BM25 gate in the extension; test guarded semantic retrieval separately |
2. The First Version Failed to Become a Habit
Collect assumed I would leave an AI conversation, open a separate local app, and organize snippets after the fact. In my own use, that extra step usually did not happen. The idea might be worth keeping, but once I had moved on, I was unlikely to return and file it.
That pointed to an entry problem rather than a missing feature. I rebuilt the workflow as a side panel beside the conversation. The user can select a passage on a supported LLM page, explicitly capture it, add a thought, and keep working. Echo stores the thought, the selected text, and source information in IndexedDB through Dexie. It is a small memory object tied to its origin, rather than a copy of an entire chat.
The scope stayed narrow. The extension supports ChatGPT, Claude, Gemini, and Grok pages. It does not create an account, sync memories to a cloud service, or decide autonomously what the user should believe. Those limits keep the first product centered on one job: making a previous thought available when the current work gives it a reason to return.
3. Designing Recall Around Interruption
The first recall probe exposed a less obvious problem. Common Chinese words such as 实用, 这个, and 快速 made unrelated records appear connected. A suggestion can be mathematically similar enough to rank and still feel like a distraction to the person reading it.
I added a Probe view that records which candidates were considered and why they passed or failed. An audit of 16 reports, 1,493 candidate scans, and 256 accepted candidates showed how much of the early behavior came from thin lexical evidence. I then tightened the gate: a single shared term is insufficient, phrases and multiple eligible terms carry more weight, common terms are downgraded, and duplicates are suppressed.
The product now treats an empty result as a valid outcome. It would rather miss some paraphrases than repeatedly interrupt the user with weak connections. That is a product choice, not a claim that BM25 understands meaning better than an embedding model.
Returning to an old thought also requires care. LLM pages can rebuild messages or load history only when scrolled. Echo's source anchors try provider message IDs, text fingerprints, nearby messages, and limited retries. If the match remains uncertain, the extension reports the failure instead of jumping to a plausible but wrong passage. Pending edits are queued through Chrome local storage so a sleeping Manifest V3 service worker does not silently lose them. A versioned JSON backup provides a way to move or recover local records.
4. Testing Whether Semantic Search Earned a Place
Lexical matching misses ideas expressed in different words, so I built an offline benchmark for BM25, raw vectors, and hybrid retrieval. I measured both useful retrieval and unwanted surfacing. The public fixture contains 55 constructed Echo records and 40 annotated queries; it reflects dogfood themes but is not an export of private conversations.
| Public fixture | Precision@3 | Recall@3 | Wrong surfaces per 100 queries | Correct abstention |
|---|---|---|---|---|
| Product BM25 gate | 28.7% | 25.8% | 5.0 | 80% |
| Raw vector top-3 | 32.5% | 76.3% | 17.5 | 0% |
| Lexical-first hybrid | 25.0% | 59.6% | 7.5 | 60% |
On this fixture, raw vectors recovered many more paraphrases, but they also surfaced unrelated results more often and never stayed silent on queries labeled for abstention. I did not move them into the extension on the strength of recall alone.
I then evaluated a later, time-sliced dogfood holdout: 21 queries over 91 unique Echo records. This set told a more complicated story. Raw vectors had fewer wrong surfaces than BM25 here (38.1 vs 47.6 per 100 queries), although neither method abstained correctly on the five abstention cases. The public-fixture claim that vectors increase interruptions does not generalize unchanged to this holdout.
| Real dogfood holdout | Precision@3 | Recall@3 | Wrong surfaces per 100 queries | Correct abstention |
|---|---|---|---|---|
| Product BM25 | 27.0% | 34.9% | 47.6 | 0% |
| Raw vector | 30.2% | 46.4% | 38.1 | 0% |
| Existing hybrid | 27.0% | 42.9% | 42.9 | 0% |
| Guarded hybrid | 42.9% | 29.0% | 0.0 | 100% |
I selected the guarded rule on a separate development window before running this holdout. It usually stays silent and surfaces at most one result. The clean numbers are encouraging, but small: zero wrong surfaces means 9/9 surfaced results were useful; 100% abstention means 5/5 abstention cases were handled correctly. The records and labels come from one user's recent work, with a model-assisted first labeling pass. They support a browser feasibility prototype, not a claim of general accuracy or readiness to replace the product path.
5. Keeping the Product Boundary Clear
Echo's core loop runs locally. The extension stores records in the current Chrome profile and does not send them to an Echo server or a model API for recall. Copying, backing up, or exporting a Probe trace is an explicit action. Local storage is still tied to the security of the browser profile and to wherever an exported file is kept, so the design includes backup and visible failure paths rather than treating “local” as a complete safety guarantee.
I also tested a separate question: whether retrieved records could support answers with citations. On a frozen 14-question public holdout, a local candidate-retrieval and generation pipeline answered or abstained correctly on 14/14, compared with 3/14 without supplied context. Human review found all 16 generated factual claims supported by their cited records. The test used a small, constructed corpus and an isolated local model run. It shows that grounded answers are worth studying; it does not put generation into Echo's contextual recall flow.
That separation matters. Seeing a related old thought is a lightweight interruption. Asking a model to synthesize an answer is a different task with different latency, privacy, and evidence requirements.
6. Where the Project Stands
The current extension keeps its precision-first BM25 recall gate. Raw vector retrieval and the earlier broad hybrid remain offline experiments. The guarded hybrid has earned a narrower next step: measure model size, browser memory, cold start, and warm query latency, then test it against a newly collected holdout before changing what users see.
Echo is still a personal dogfood and alpha project. I do not have evidence of retention across a broad user base, and the holdout is too small to claim stable performance. The concrete outcome so far is a product workflow that I use, an evaluation trail for its recall decisions, and a decision to leave promising but unproven retrieval methods out of the user-facing loop.
The lesson I would carry into the next project is simple: retrieval quality is not just whether the right item exists in the top three. In a tool that lives beside someone else's work, it is also whether the suggestion deserves their attention now.
Evidence: Precision decision · Public benchmark · Dogfood holdout and limits · Grounded-generation holdout · Privacy and threat model

