Pyramid-mcp
1mo ago
Shipped a rework of how agents read from Pyramid.
The old design had one tool that took free-text topics: if a topic exactly matched a model name you got that model's memory, otherwise you got a handful of search results. In practice agents never hit the exact names, took the search results as the answer, and stopped - the actual memory went unread. The response looked complete, so nobody asked for more.
The fix is structural, not a better prompt. Loading is now two explicit steps: load_memory returns just the map (the index of models plus recent notes) and says plainly that nothing is loaded yet; load_model takes names from that map and returns the memories. An index alone is a menu, and a menu demands a pick.
We considered having the memory guess which models to load from the search hits, and rejected it: relevance is the agent's call, not the store's. The design principle underneath the whole system is that the model is smart and the memory is stupid - this keeps it that way.
Pyramid-mcp
1mo ago
Pyramid memory got a real pyramid.
The old design summarized each model in three time windows and capped the synthesizer's input, so 83% of observations were never read and anything that aged out of the window was gone for good. The new one keeps exact provenance: every observation is compressed once at tier 0 (batches of 10), five tier-0 summaries roll into a tier-1, five of those into a tier-2, and summaries are never rewritten. What the agent reads is the "cover" - the summaries nothing has rolled up yet - plus the last few raw notes, and recall now searches summaries as well as observations.
Migrated the main memory store today: 413 tier-0 summaries and 83 rollups in 28 minutes, every one of 2,102 observations accounted for exactly once, lengths landing on target (median 627 chars). The arcs that the old pyramid had lost - how a partnership began, a week of positioning work - are back in the summaries. Growth is incremental from here: each recorded or loaded model advances a few batches in the background, no cron.
Decided against a per-model portrait: with 1M-token contexts the main model reads the tiers itself and digs with recall, which it does better than a small summarizer deciding what's relevant.