Branch, not yet merged: steps 04–18 below trace makeup_search.py on experiment/two-step-classify-section (repo: code.straive.com/tpavan.kumar/straive-slide-search) — pushed, MR open, not yet in main / not yet what production actually serves today. Every trace number is still real: run against the real 122,623-slide corpus through the real local engine code on that branch, not simulated or invented — just not live for end users yet. Steps 01–03 and 18b/END (chat app + explain endpoint) are unaffected by this branch and remain accurate to production. NEW-tagged blocks were added/changed since this page was first built — everything else was already live before that.
Even more local, uncommitted, on a different branch (fix/claim-prompt-vague-language): steps 04b, 06b, and 14b below are real, direct-tested fixes built this week from a self-improving eval loop against real sales-team queries — none of the three are committed or pushed yet, so they have no git version-history panel (an honest gap, not an oversight). Every number in them is from a real, direct test on the live 122,623-slide corpus, same standard as every other block on this page.
Every dashed EXAMPLE box below is real: run against the live 122,623-slide production index, not invented. One query — "customer analytics" — is traced end to end through every stage so the numbers stay consistent.
New: click ▸ version history on any block to see exactly when it was created, why, and every real commit that has touched it in the last 3 months — pulled straight from git, not summarized.
The new pipeline — 23 stages
Same shape as the old 11-stage summary, updated for this branch. NEW = added or restructured in this redesign; the rest were already live before it. Full detail, real examples, and version history for each are in the numbered cards below.
01. Tool routingChat agent now defaults to search_v2 for every brief, simple or multi-part.
02. Backend handoffMCP call reaches the real search engine — no filtering or ranking yet.
03. Multi-angle split NEWA multi-part brief runs each topic through classify+expand independently, then merges.
04. Classify + two-step scopeOne merged LLM call splits "what to SEARCH" (search_types) from "what to SHOW" (answer_sections).
04b. Explicit exclusion carve-out NEWDeterministic regex: "without case studies" actually excludes them, overriding the step 04 safety net for that one query.
05. Synonym/cluster expansionCurated + taxonomy + AI-suggested terms, capped at 10, real similarity cutoff (≥0.5) on AI terms.
06. Hard candidate prefilter NEWReal search pool built from search_types BEFORE any semantic search — genuine hybrid, not post-hoc.
06b. Named entity / phrase / person detection NEWThree rarity-based, deterministic checks widen the pool and boost score for anything narrow and specific.
07a. Primary RRFExact semantic search of your own words, within the prefiltered pool only.
07b. Boost RRFSame, for the ≤10 expansion terms — a bounded, capped assist, never an equal vote.
08. Folder × type × recencyMultiplies by three real internal signals to get each slide's final score.
09. Exact-content dedup NEWDrops verbatim-duplicate slide text copy-pasted across different decks.
11. Global per-deck cap NEWCaps one deck to 2 appearances across the WHOLE response, not just per section.
12. Relevance floor NEWDrops anything under 15% of this query's own single best score, across every section.
13. Group into 8 sections NEW2 new categories carved out; verified rows append after real Case Studies, never displacing them.
14. Section (15) + global (50) cap NEWHard backstops, section cap applied first, then the whole-response cap.
14b. Guaranteed entity slot NEWAfter caps, splices in one real match for a named entity/phrase/person if nothing else survived.
15. Sections ordered by relevance NEWSorted by each section's own top score for this query, not a fixed sequence.
16. Coverage check NEWFlags a genuine "nothing in the corpus" gap vs. a ranking miss.
17. Response assembled NEWVerified rows now flagged inline per-card; the old master_sheet list is always empty.
18. Agent surfaces resultspresent_groups renders cards in the backend's own order — no re-splitting merged sections.
18b. "Why this?" explain NEWOn-demand, per-card match explanation, batched across visible cards.
How much gets filtered out, in one picture
Real trace, the same "customer analytics" query threaded through every example on this page. Each bar is one checkpoint in the pipeline above.
Whole corpus before anything runs
122,623 slides
Step 06 hard prefilter by search type
97,262 eligible
− 20.5% excluded, before search even runs
Steps 08–11 ranked, scored, deduped, grouped
60 candidates
15 slots × 4 relevant sections
Step 12 relevance floor
21 candidates
− 39 too weak vs. the query's own best result
Steps 13–17 final response
34 cards shown
+ 13 curated case studies blended back in
Bars are log-scaled, not linear — on a straight linear scale, every bar after step 06 would be too thin to see next to 122,623. The small uptick at the very end isn't a mistake: that's step 13 blending curated, verified case-study content back in after the floor already trimmed the real matches down.
STARTUser types a query in the chat app
e.g. "customer analytics", or a multi-part brief like "fraud detection use case, regulatory compliance automation" · Agentic Chat App
in plain English
A query is simply whatever you type into the chat box — one topic, or several combined in one message (a "multi-part brief"). Everything on the rest of this page traces that one piece of text all the way through to the results you see.
the logic here
You type a question or brief into the chat box.
Nothing gets judged or decided yet at this point — this step just captures exactly what you asked, word for word, to hand to the next step.
running example
Query: "customer analytics"
01Chat agent decides which tool to calllib/agent/gemini.ts
System prompt now positions search_v2 as the primary tool for every brief, simple or multi-part — it used to lose out to the plain search tool for multi-outcome briefs specifically because it couldn't take a list. Universe scope (DAAIS/S&R/EdTech) is injected automatically into whichever tool actually gets called.
in plain English
The system prompt is simply the written instructions the assistant's AI is given before every conversation, telling it how to behave and which tools it's allowed to use. search_v2 and plain search are just the names of the two different search tools it can pick between — this step is the assistant deciding which one to actually use for your question. Universe scope is the DAAIS/Sales & Research/EdTech restriction you can set from the toggle at the top of the chat, if you've picked one.
the logic here
The assistant reads your message and its own written instructions.
Because those instructions now say search_v2 is the go-to tool for every kind of request, it picks that one rather than the older, simpler search tool — regardless of whether your question is one topic or several.
If you'd set a universe filter on the toggle, it gets automatically attached to whichever tool actually gets called, so it can't be silently dropped.
Fixed: universe scope used to be silently dropped whenever the model called plain search instead of search_v2 — enforcement now covers both tool names.
02MCP tool call reaches the backendslide_search_server.py
search_v2(query, universes, top_per_group) → makeup_search.get_makeup_engine().search(...). query can now be a single string or a list of outcome angles.
in plain English
This is the moment your search physically leaves the chat app and lands on the actual server holding all 122,000+ slides. The code shown is just that handoff: your search text (query), your universe filter (universes), and how many results to return per section (top_per_group) get handed to the real search function. An "outcome angle" is simply one topic within a multi-part brief — see the next step for how several of them get handled together.
the logic here
Take the tool call the assistant made in step 01, and route it to the actual backend server.
Call the one real search function everything else on this page runs through, handing it your search text, your universe filter, and how many results you want per section.
No filtering or ranking decisions happen yet — this step is purely the handoff.
03Multi-angle normalizationsearch()NEW
A single query becomes a one-item list; a real multi-part brief keeps all its angles. Every angle runs intent classification + synonym expansion, and the results merge (intents unioned, expansion terms deduped case-insensitively).
in plain EnglishIntent classification (next step) figures out what KIND of result each topic wants; synonym expansion (two steps after that) thinks of related terms for it. "Unioned" just means combined into one list with no duplicates; "deduped case-insensitively" means treating "CDP" and "cdp" as the same term when clearing out repeats. All of this happens once per topic, then gets merged together at the end.
the logic here
Check whether the question is one topic or genuinely several.
If one, wrap it as a list containing just that single topic. If several, keep each as its own separate item in the list.
Send every item in that list, independently, through its own full round of steps 04 and 05 below — no topic influences another topic's judgment.
Once every topic has its own result-type judgment and its own related terms, merge everything back into one combined judgment and one combined, duplicate-free list of related terms.
Replaces the old separate classify step — one merged LLM call (still gemini-2.5-flash) now does three jobs at once: ranking intents (as before), plus TWO distinct section lists instead of one. search_types decides what gets SEARCHED (generous — "could this type plausibly hold a good answer?"). answer_sections decides what gets its own HEADING in the reply (selective — "would a person asking this exact query expect to see this as a section?"). Every entry in answer_sections must also be in search_types, but not everything searched earns a heading.
in plain English
Before this change, one question ("which sections matter here?") did double duty as both "what should I search" and "what should I show" — so a query could never search a content type it wasn't also going to display a heading for. Now those are two separate answers from the same AI call: search_types is the wider net (cast broadly, so nothing good gets missed), answer_sections is the narrower "what's actually worth its own heading" — a type can be searched purely to widen the net without ever becoming a visible section.
the logic here
Take one topic and send it, with the fixed list of content types (now eight, not six — see step 13) and worked real examples, to the AI model.
Step A: the model picks every type that could plausibly hold a good answer — inclusive, errs toward including.
Step B: of those, the model picks which should actually get their own heading — selective, only what this exact query would expect.
Separately, the same call still picks ranking intent(s) (case_study / approach / capability / collateral / demo) exactly as before, defaulting to one unless there's an explicit second signal.
If this AI step fails outright, both lists fail OPEN (all eight types) rather than risk an empty response, and intent falls back to the same keyword heuristic as before.
example — real trace, "customer analytics"
intents = ["capability"] · search_types = answer_sections = ["Data & Analytics", "Capabilities & Overview", "Approach & Methodology", "Case Studies"] (identical here — a query where they genuinely diverge: "fraud detection case study banking" searches Data & Analytics but doesn't give it a heading, since the query only asks for case studies)
Bug found + fixed in this same change: the always-force-Case-Studies-into-both-lists rule (kept from the older hard-section-filter fix) means Case Studies is the one type that doesn't get the new "searched but not headed" treatment — a deliberate, disclosed trade-off, not an oversight.
A deterministic regex, not another AI call, checks the query for an explicit negation phrase — "without / excluding / exclude / no / not including / skip(ping) case studies." If it matches, "Case Studies" is removed from BOTH step 04 lists for this query only, and the step 04 always-include safety net is skipped for this one query so it can't silently add Case Studies back.
in plain English
Step 04 has a built-in safety net that always puts Case Studies back into the results, even if the AI classifier didn't pick it — built for a real earlier bug where a topic-only query missed the single best case-study match in the whole corpus. But that safety net had no exception for someone explicitly asking to LEAVE case studies out. This step is that exception: a plain phrase check, not another AI judgment call, so it can't misfire the way an LLM guess could.
the logic here
Check the raw query text against a fixed pattern for phrases like "without case studies" or "no case studies."
If it matches, remove "Case Studies" from both this query's search_types and answer_sections.
When step 04's own always-include safety net runs, skip re-adding Case Studies for this one query only — every other query keeps the exact same guarantee as before.
Real bug this closes: a real logged sales query — "create a second deck without case studies and demoes... only retain the capability slides" — still returned Case Studies cards anyway, because the always-include safety net had no carve-out for explicit exclusion. Traced directly to 31 of 818 "wrong" verdicts in one measurement pass coming from just 2 real queries shaped like this one. Verified after the fix: the same query now returns 0 Case Studies cards; an ordinary case-studies query is unaffected.
✕ before
Query: "...deck without case studies and demoes." Response still includes Case Studies cards — every one graded "wrong," since the user explicitly asked for none.
✓ after
Same query. 0 Case Studies cards in the response. Every other section, and every query that doesn't use exclusion language, is completely unaffected.
Curated cluster dictionary: 96 keys / 164 terms, mined from the real corpus. Taxonomy-nearest: threshold-driven (cosine ≥0.68), up to 300 terms. World-expansion (LLM, only fires if no curated cluster matched): now filtered by a REAL embedding-similarity cutoff against the query (≥0.5) before being eligible at all — calibrated against real term/query pairs, not guessed. All three sources merge, dedupe, and cap at 10 expansion terms (raised from an earlier 4-term cap per direct instruction: "don't restrict to 4, keep 10 as the upper limit, but have a strong cut-off for inclusion").
in plain English
Same 0-to-1 cosine similarity score explained in step 10 later on this page. The 0.5 cutoff on the AI-generated terms is the "strong cut-off for inclusion" — real testing showed genuine synonyms score 0.55–0.90 against the query, while loose/tangential ones the AI suggested anyway score 0.33–0.45, so 0.5 sits cleanly in the real gap between them. Raising the ceiling to 10 does NOT mean always returning 10 — it means "up to 10, if they clear the bar."
the logic here
Check the query against the 100-cluster dictionary below — any exact trigger match adds every term in that cluster.
Separately, compare the query against ~5,600 taxonomy terms, keeping everything ≥0.68 similarity.
If no curated cluster fired, also ask the AI model for related terms — then discard any of ITS OWN suggestions that don't independently clear 0.5 similarity to the query, a real check the AI's own "prefer fewer" instruction can't be trusted to enforce alone.
Merge all three sources, dedupe, and keep the query itself plus up to 10 more.
example — real trace, "customer analytics"
10 terms kept: Unified Fan 360, Fan 360, Golden Customer Record, Single Customer View, Identity Resolution, Identity Matching, Customer Data Platform, CDP, Behavioral Segmentation, Cohort Analysis
06Hard candidate-pool prefilter, built from search_typessearch() → allowed_idsNEW
Before ANY semantic search runs, the real candidate universe gets built: every slide whose own section is in search_types from step 04 (the wide list, not the narrow display list) AND whose universe (DAAIS/S&R/EdTech) matches, if scoped. Real vectors for just that subset get reconstructed once, reused for both the primary and secondary search passes below — not the approximate top-300 search over the whole 122k index this page used to describe.
in plain English
This is the literal answer to "you asked for X, why did I get Y" complaints: previously, everything got searched first and irrelevant types got softly down-ranked afterward — a strong-scoring slide from a type nobody asked for could still sneak in. Now that type is excluded from the search entirely, before ranking even starts, if step 04 didn't put it in search_types.
the logic here
Take the search_types list from step 04 (all eight are eligible if nothing was excluded).
Keep only slides whose own content-type section is in that list, and whose universe (if the user scoped one) matches.
Reconstruct the real, exact vectors for just that subset — a genuine brute-force search within it, not an approximate shortcut over everything.
example — real trace, "customer analytics"
candidate pool: 97,262 of 122,623 slides eligible (79.3%) — everything tagged Data & Analytics, Capabilities & Overview, Approach & Methodology, or Case Studies; Team & Credentials, Credibility & Recognition, Commercial & Contract Terms, and Other were excluded before search even ran.
06bIs the query naming something specific? Three rarity-based detectors_query_entity_terms() · _detect_person_name_terms() · _query_phrase_terms()NEW
Three independent, deterministic checks on the query text — none call an LLM — feed the SAME downstream mechanism: (1) any single word appearing in ≤300 of the corpus's 122,623 slides (≤0.24%), casing-independent; (2) only if the query itself contains "profile / resume / bio / who is," any capitalized name span in it, skipping the very first word since sentence-initial capitalization is just grammar; (3) 2–3 word adjacent phrases (after generic-word filtering) checked against the same ≤300-mention rarity threshold. Anything caught gets unioned into step 06's candidate pool even outside search_types, gets a 1.5× score boost wherever it appears (step 08), and gets a shot at a guaranteed slot if nothing else survives (step 14b).
in plain English
This is the system asking itself "did they name something narrow and specific — a client, a person, a named product?" A word showing up in only a couple hundred of 122,623 slides almost certainly IS the specific thing being asked about, not a generic business word. The person-name check runs separately because a real name like "Naveen Gattu" can appear in hundreds of slides as routine leadership-team boilerplate — rarity alone wouldn't flag it, so this one instead watches for query PHRASING ("who is...", "...profile") rather than word frequency. The phrase check exists because some asks are narrow as a PAIR of words even though neither word alone is rare — "capital call reporting" has no individually rare word, but "capital call" together appears in only 50 real slides.
the logic here
Pull the query's distinctive words (skipping the same generic-term list used everywhere else on this page).
Word check: look up each one's real corpus mention count (cached); keep it if that count is 1–300.
Name check: only if the query itself asks for a profile/bio/who-is, scan for capitalized word spans, skipping the query's very first word.
Phrase check: build every 2- and 3-word adjacent span from the filtered words, look up each one's real corpus mention count the same way, keep it if 1–300.
Union everything all three catch into one set for this query, passed forward into steps 06, 08, and 14b.
example — real test, "capital call reporting"
No individually-rare word (capital/call/reporting are all common alone) — but the phrase check finds "capital call" at only 50 real corpus mentions. Before this step: buried. After: "Private Equity Capital Call Notice Automation" surfaces as the real #1 result.
example — real test, "Anand S profile" / "Naveen Gattu profile"
"Anand S": name check fires, real matches jump straight to the top of Team & Credentials. "Naveen Gattu": name check also fires and does surface 4 genuine matches (up from zero) — but he's a well-known co-founder named in 382 different real slides (routine org-chart/leadership boilerplate), so a fixed 1.5× boost alone doesn't discriminate well inside that large a pool. Reported honestly here as a real, narrower limitation, not hidden.
07aPRIMARY: exact semantic search, within the prefiltered pool_rrf_pass_subset(queries)
Each angle is embedded and matched against the REAL vectors for the step 06 candidate pool only — an exact brute-force cosine search within that subset, not an approximate ANN search over the whole index — fused via Reciprocal Rank Fusion, every angle an equally-weighted voter.
in plain English
This step searches using your own actual words — none of the related/synonym terms from step 05 are involved yet, and only slides step 06 already let through are even candidates. Every slide gets a primary number: the higher it ranked in this word-for-word search, the bigger the number. A slide that didn't show up at all gets exactly 0.
the logic here
Take your own exact wording — none of step 05's related terms are used yet.
Search the step 06 candidate pool's real vectors with it (not the whole corpus).
Every slide that turns up gets a primary score based on how high it ranked.
Anything that doesn't turn up gets primary = 0, carried into 07b.
example — RRF formula
Each hit contributes 1 / (RRF_K + rank), RRF_K=60, summed per angle. A slide ranked #3 contributes 1/(60+3) = 0.0159.
07bSECONDARY: bounded cluster-term boost, same pool_rrf_pass_subset(boost_terms)
Every expansion term from step 05 (up to 10, all real) searched against the SAME prefiltered pool, pooled into ONE normalized signal (0–1), capped at +35% lift on top of primary. A term with zero primary relevance gets a heavily discounted score, never a free ride to the top.
in plain English — what "primary" and "boost" mean hereprimary is the score from step 07a — 0.0 means it didn't come up at all on your literal words. boost is built from the expansion terms — the more of those a slide touches, the higher its boost. If a slide already has a real primary score, related terms lift it by up to 35% more. If primary is a flat 0, related terms alone can only add a small, heavily discounted amount — never enough to push a genuinely unrelated slide to the top by themselves.
Tried and reverted in this same change: weighting each expansion term's vote by its own similarity to the query (so a barely-relevant 10th term counts less) was implemented, tested on a real full re-run, and made ZERO measurable difference — reverted rather than keep unproven complexity that costs an extra embedding call per query.
Score is multiplied by (a) how strongly the slide's folder matches the query's intent — a purely internal ranking signal, never shown to the user — (b) whether the slide's own content type matches the intent, and (c) a mild recency factor.
in plain English — what each number below means
By this point every slide already has one combined score from the two steps above, called the fusion score. This step multiplies that score by three separate adjustments: folder_weight rewards a slide for living in a folder that fits what the search is looking for; type_weight rewards or penalizes a slide based on what KIND of slide it actually is; recency_weight gives a small nudge to more recently created content. Multiplying these together means a slide has to do reasonably well on all three to end up near the top.
the logic here
Take the combined fusion score a slide already earned from steps 07a and 07b.
Look up which folder the slide's deck lives in, and multiply by how well that folder type matches the intent from step 04.
Look up the slide's own content-type tag, and multiply again — boosting a genuine content match, or cutting the score by more than half for a known low-content type (unless the content-rich safety-net exemption applies).
Multiply once more by a small recency factor based on when the deck was created, favoring newer content slightly.
Whatever comes out of multiplying all three adjustments together is the slide's real, final score — used for every step from here on.
example — full final score, real slide, real today's trace
"Identified target customer groups for rolling out efficient Cross-sell/Up-sell marketing initiatives" — deck Hi-Tech_and_Consulting_20251012.pptx, folder /DAAIS/Sales/Collateral/Technology Capabilities, type dashboard (real content: NPS distribution, revenue per account, customer journey contributions — a genuine BI artifact, not just a narrative slide that mentions data).
fusion score (from 07a+07b) = 0.0202
folder_weight (capability × Collateral) = 3.0
type_weight (capability × dashboard) = 1.3(genuine data-type boost)
recency_weight = 1.244
FINAL = 0.0202 × 3.0 × 1.3 × 1.244 = 0.0978 ← real #1 in Data & Analytics for this query today
raw fusion score
0.0202
× folder weight (3.0)
0.0606
× type weight (1.3)
0.0788
× recency (1.244) = final
0.0978
Corrected after a stakeholder question caught it: the previous version of this example named a different, same-titled slide from a different deck (real content dedup/title collision) — the numbers were internally consistent but for the WRONG slide. This is the one the trace actually returned.
Since this week, a 4th multiplier applies:entity_boost — 1.5× if the slide's own title/snippet literally contains a term step 06b flagged as a rare entity, phrase, or person name, else 1.0× (no effect). Deliberately capped at 1.5×, not higher — real testing showed a bigger multiplier risks a "rare" term dominating ranking on queries where that term isn't actually the point of the question; see step 14b for the separate guarantee mechanism built for the cases where 1.5× alone still isn't enough.
09Exact-content dedup_dedupe_content()NEW
Drops lower-scored results whose real slide text is an exact (normalized) match to a higher-scored one already kept — catches a boilerplate slide copy-pasted verbatim into different decks, which deck-title dedup can't see at all.
in plain English"Normalized" just means minor formatting differences — extra spaces, capitalization — are ignored before comparing two slides' text. So two slides count as an exact match if their real content is identical, even if one happens to be spaced slightly differently. "Dedup" is short for de-duplication: removing repeats so you don't see the same real content twice.
the logic here
Sort every remaining candidate slide by its final score from step 08, highest first.
Walk down that sorted list one slide at a time.
For each slide, check whether its real (normalized) text has already been seen, word for word, from a higher-scoring slide earlier in the list.
If it has, drop this lower-scoring copy entirely. If not, keep it and remember its text, so later, lower-scoring duplicates of it get caught too.
example — real pair caught
...nfl-final-submission_Slide_50 and ...nfl-unified-fan-view-v1_Slide_50 — two different deck files, snippet text word-for-word identical ("5. Data & Analytics Partner for leading Technology Major…"). Lower-scored one dropped.
Catches the harder case: the same template slide reused across decks with differing extracted text (different vision captions/wrapper format per file), so exact-text matching misses it. Reconstructs each candidate's FAISS vector, computes the full pairwise similarity matrix, and unions any pair ≥0.90 cosine via union-find — not a greedy "compare only to what's already kept" walk, which left survivors when one boilerplate slide had 254 near-identical copies smeared across a 0.74–0.99 similarity range.
in plain Englishcosine similarity is just a similarity score from 0 to 1 for how close two slides are in meaning — 0 means nothing alike, 1 means essentially identical content. This step checks that score for every possible pair of candidate slides, then uses a technique called union-find to correctly group together every slide that's similar to at least one other slide in the same chain — even if not every pair in that chain scores high enough directly against each other (see the real example below, where A-to-B and B-to-C are both similar enough even though A-to-C isn't). A simpler "only compare to what I've already kept" approach misses exactly this kind of chain, which is why it's not used here.
the logic here
Take every slide still in the running after step 09.
Compute how similar every single pair of them is in meaning (not just exact text), using the same real math the search index already relies on.
For every pair that scores 0.90 or higher, mark them as connected.
Follow all the chains of connections — if A connects to B and B connects to C, treat A, B, and C as one group — so a smeared-out set of near-copies all end up together even if not every pair individually cleared 0.90.
From each group, keep only the highest-scoring slide and drop the rest.
example — real 15-instance sample, one boilerplate template
pairwise cosine similarity: min 0.74, max 0.99, mean 0.91
28 of 105 pairs individually score BELOW the 0.90 threshold
→ a greedy walk (compare only to what's kept) leaves 3 survivors→ union-find (compare via ANY chain of similar pairs) collapses to1 survivor
(13 of 15 instances merge into one connected component)
Caps one canonical deck to 2 appearances across the whole response, applied before the group split. The existing per-group DECK_CAP=2 (below) couldn't stop a deck from claiming 2 slides in every section it touched — a large compilation deck could still show up 4× total.
in plain English
A "canonical deck" just means different draft versions of the same underlying deck (V1, V2, vFinal, and so on) are recognized as one single deck for capping purposes, not counted separately — explained with real numbers in step 13's version history below. The "group split" is simply the moment results get divided up into the eight sections described in step 13.
the logic here
Take the full, deduplicated, score-sorted list from step 10.
Work out each slide's real underlying deck identity, with version/draft naming stripped off, so different drafts of the same deck count as one.
Walk down the list keeping a running count of appearances per deck.
The moment a deck would be about to appear for a 3rd time anywhere in the whole response, skip that slide and move to the next one — capping every deck to at most 2 total appearances before anything gets split into sections.
example — real, live-observed before this fix
"Visualisation-slides Consolidated.pptx" (a large multi-hundred-slide compilation deck) appeared 2× in Case Studies + 2× in Proposals = 4× total for one query. After the fix: capped to 2× total, not eliminated — the other 2 slots go to genuinely different decks.
Drops anything scoring below 15% of THIS QUERY's own single best result, across every section — not each section's own top (a per-section-relative floor does nothing for a whole section that's uniformly weak compared to the rest of the response).
in plain English
Before this, top_per_group=15 just meant "fill up to 15 slots per section" with no floor on how weak a result could be and still take a slot. This step asks a different question: compared to the single BEST result anywhere in this whole response, is this candidate even in the same ballpark (within 15%)? If not, it gets cut rather than padding out a section that's genuinely thin for this query.
the logic here
Find the single highest score anywhere across every section for this query.
Multiply it by 0.15 to get the floor value.
Drop every remaining candidate scoring below that floor, in every section, not just the weak ones.
13Group into 8 real content types; Case Studies + Verified merge into onecontent_group_of() · _diversify() · _blend_verified_into_case_studies()NEW
Sections are now eight, not six: Case Studies / Approach & Methodology / Data & Analytics / Capabilities & Overview / Team & Credentials / Credibility & Recognition / Commercial & Contract Terms / Other — the last two carved out of the old catch-all "Capabilities & Overview" bucket (91.7% of it carried a single generic upstream tag, hiding real, distinct populations: awards/analyst-recognition slides and scope-of-work/pricing slides, found by sampling real titles and confirmed at corpus scale, ~1,570 slides total). Verified Case Studies is no longer its own section — the team's hand-curated master-sheet rows now merge directly into "Case Studies."
why eight, specifically — not a round number, two evidence-driven stagesStage 1 (the original six): sections used to come from which Drive FOLDER a slide lived in (client/domain-cap/tech-cap/collateral) — but a folder conflates "who's this deck for" with "what's actually on this slide," so it was rebuilt from each slide's own real content-type tag instead, counted directly across all 122,623 slides (content: 47,341 · diagram: 14,112 · case_study: 13,223 · table: 6,649 · framework: 2,922 · dashboard: 1,272 · team: 1,546, etc.) and grouped into six sections describing what's actually ON a slide.
Stage 2 (the two new ones): added on a direct stakeholder request to expand the taxonomy further "if useful." Checking found one of the six, "Capabilities & Overview," was so generic it was silently burying two things reps actually ask for by name — award/analyst-recognition slides and scope-of-work/commercial-terms slides. Sampling real titles found exactly 872 and 644 real matches for those two — common and distinct enough to earn their own heading, unlike anything else sampled.
So it's six from what the corpus actually contains, plus exactly two more that real sampling proved were being hidden — not a target headcount. A ninth would need the same evidence: a genuinely common, distinct population currently invisible inside an existing bucket.
in plain English — what "gets a heading" means here
The eight section types described here decide what gets its OWN HEADING (from step 04's answer_sections) — but note that's a different question from what got SEARCHED (step 06's search_types). A type can be searched to widen the net without ever becoming a visible section, and a section only appears here at all if step 04 said it should.
the logic here — grouping
For each surviving slide, look up its content-type tag and decide which of the eight sections it belongs in — including two new regex-matched carve-outs (award/Gartner/"why choose us" language → Credibility & Recognition; scope-of-work/commercial-terms/payment-terms language → Commercial & Contract Terms), computed once when the engine loads, not per query.
Within each section, cap any one deck to 2 slides.
the logic here — merging verified case studies in
Separately search the team's 1,194-row master case-study sheet by cosine similarity (same mechanism as always — see the merged content below).
Real Case-Studies slides keep their existing, already-trusted ranking order UNTOUCHED — never displaced.
Verified rows append AFTER all of them, each given a synthetic score strictly below the weakest real slide's, ordered only among themselves by their own similarity.
Two real bugs found and fixed building this: (1) a scale-mismatch bug let verified-only sections use a RAW cosine score (0.5–0.9) directly as display score, wrongly outranking every other section's RRF score (0.01–0.15) — a Johnson & Johnson PowerBI-team card once won "leadership team credentials" purely on scale. (2) an earlier design let the single best verified row always compete for the #1 OVERALL slot — live testing found master-sheet similarity doesn't reliably separate strong from weak matches (a weak match scored in the same band as strong ones), so this regularly promoted weak verified rows above strong real ones. Both fixed: verified rows never displace a real match now, confirmed on a real 100-query before/after re-run.
example — real Case Studies blend, "customer analytics" today
2 real slides survive the relevance floor (0.0154, 0.0151)
13 verified rows append after, decaying (0.0150 → 0.0141), all below the weakest real one
→ 15 total, 2 real + 13 verified, never crossing over each other
the logic here — quantified-claim gate (2026-09-25, built + merged separately from the experiment/two-step-classify-section branch above — already in main)
Separately from step 04's classify call — a genuinely distinct, second LLM call, deliberately NOT folded into that shared one — the query gets checked for whether it EXPLICITLY demands a specific kind of quantified proof: dollar figure, percent, or FTE/headcount.
If it demanded at least one, every Case Studies row (real slide AND verified-sheet row alike) gets checked with a plain, deterministic regex against its OWN real title+snippet text — not another AI guess — for whether it actually states a matching dollar/percent/FTE figure.
Any Case Studies row that doesn't state a matching figure is dropped from THIS SECTION ONLY — Approach & Methodology and Capabilities & Overview are untouched, since discussing ROI conceptually without a hard number yet is still legitimate content there.
If the query didn't demand a specific proof type, this gate is a complete no-op — behavior is unchanged.
Real bug this closes: a query demanding "quantified ROI proof points... dollar savings, FTE reduction" was matching a case-study slide whose only real number was "1.6ktCO₂ emissions reduction by cloud migration" — topically about cost/efficiency, but not the TYPE of number actually asked for. The old ranking had no check for which type of number a slide's own text states, only whether it was topically in the neighborhood.
example — real poison case, before vs after
Query: "...quantified ROI proof points... dollar savings, FTE reduction." Before: top Case Studies match stated only a carbon-emissions figure. After: that slide is excluded from this section — the gate requires an actual dollar/percent/FTE figure in the slide's own text before it can rank here.
Verified on 100 real historical queries, re-run end to end:
gate activated on 2/100 — for both, every result it would have removed had already
independently failed to rank near the top anyway, so the top result was unchanged either way
→ 0 regressions across all 100 (verdict-class comparison, before vs after)
✕ before
Query: "...quantified ROI proof points... dollar savings, FTE reduction." Top Case Studies match: a cloud-migration slide whose only real number is 1.6ktCO₂ emissions reduction — a carbon figure, not the dollar/FTE figure actually demanded.
✓ after
That slide is excluded from Case Studies for this query. Only slides whose own text states a real dollar, percent, or FTE figure can rank in this section when the query explicitly demands one.
you'll get asked this
Why these 8 categories specifically — why not generate new categories dynamically, curated to each query?
Because these 8 aren't picked per query — they're a property of the SLIDE, decided once, corpus-wide, when the engine loads, not something invented fresh for whatever you happen to type. The six original ones came from literally counting what content types exist across all 122,000+ slides and grouping those into six buckets that describe what's actually on a slide. The two we added later weren't invented either — we found one of the six, Capabilities & Overview, was so generic it was quietly hiding two things reps ask for by name, and we only added those after sampling real titles and confirming 872 and 644 genuine matches. So the bar for a new category is real evidence of a common, distinct population being hidden — not "this particular query would like its own bucket."
If we let categories get invented per query instead, a few things break. First, the same slide could get a different heading every time depending on how someone phrased the question — that destroys consistency, and it destroys the sales team's ability to build muscle memory around "Case Studies is always where the proof points live." Second, every piece of tuning we've built — the case-studies-always-shown safety net, the dollar/percent/FTE claim filter, the guaranteed-entity-slot mechanism, the verified case-study blend — is anchored to these exact section names. If the names are fluid, none of that logic has a stable target to attach to. Third, we track "wrong rate" per section over time as part of how we're improving this thing — that kind of measurement only works because the section is the same thing query over query, not a new ad hoc bucket every time.
And the real need behind "shouldn't specific queries get something tailored" is already handled — just not by inventing new headings. That's what the entity, phrase, and person-name detection (steps 06b/14b) does — when a query names something narrow and specific, we don't create a new category for it, we make sure the right slide surfaces inside the EXISTING categories, boosted and guaranteed a spot if it needs to be. The specificity is handled at the ranking level, inside a stable structure, not by making the structure itself unstable.
14Per-section cap (15) + whole-response cap (50)_apply_global_cap() · SECTION_RESULT_CAP=15 · GLOBAL_RESULT_CAP=50NEW
A hard per-section ceiling of 15 — independent of whatever top_per_group a caller asked for — applied before the existing whole-response cap of 50 across every section combined.
in plain English
Two backstops, applied in order: first, no single section can ever exceed 15 cards, however many survived the floor above. Then, across the WHOLE response, no more than 50 cards total — a query with several genuinely strong sections still can't overwhelm a rep with more than they'll actually look at.
the logic here
Truncate each section to its own top 15 (already sorted best-first, so this keeps the strongest).
If the total across all sections still exceeds 50, keep the globally highest-scoring 50 regardless of section — a strong section keeps more slots, a weak one keeps fewer.
example — real trace, "customer analytics"
21 candidates total after the floor → all under both caps, none dropped here
(the 15-cap DOES bind on the merged Case Studies section above: 2 real + 13 verified = 15 exactly)
14bGuaranteed slot for a named entity/phrase/person, if nothing else survived_ensure_entity_representation() · _literal_entity_ids()NEW
Runs AFTER the section and global caps above, not before. If step 06b detected a rare entity/phrase/name AND none of its real, literal matches survived natural ranking into the final response, splice in its single best-scoring real match into the top answer section. Capped at exactly 1 extra slot — a floor for real content the user explicitly named, not a takeover of the section.
in plain English — why this exists on top of the 1.5× boost
Real testing on "blackrock case studies" showed the boost alone wasn't enough: every real BlackRock mention in the corpus is classified outside Case Studies entirely (Capabilities & Overview / Additional Materials), so step 06's pool-widening was needed just to make it eligible at all — and even then, a literally-matching but generic-looking slide (an overview page, not case-study-formatted content) still sat too far below genuinely case-study-shaped content on raw score for a 1.5× multiplier to close the gap. This step is the deliberate backstop: guarantee ONE real, literal match survives, rather than raise the multiplier further and risk it distorting unrelated queries.
the logic here
Only runs if step 06b found something and the query has at least one display section.
Check the top display section: did any of step 06b's real matches make it in naturally?
If yes, do nothing — the guarantee is a backstop, not an override.
If no, take the single best-scoring real match among step 06b's candidates and splice it into that section, re-sorted by score.
Placement bug found and fixed while building this: this call was first placed right after grouping, BEFORE the floor/cap stages above ran — direct inspection of the actual output JSON showed the spliced-in item got cut straight back out by the diversify step's own top-N selection before floor/cap even had a chance to run, so the guarantee had zero real effect. Moving the call to run on the FINAL groups, after every other trimming stage, fixed it — confirmed by inspecting the resulting response directly, not by re-reading the code.
A stricter version was tried and reverted (real incident): a hard exclusion filter — drop any Case Studies card that doesn't literally name the detected entity — was tried first instead of a soft guarantee. Real testing showed it turned "blackrock case studies," "AI Data Modeler case studies," and "case studies on Salesforce Harmonization" into zero-result queries, because the real matching content for those entities is classified entirely outside Case Studies, so nothing survived the hard filter at all. Standing rule this violates: zero results is the worst possible outcome for a sales rep — worse than a tangential result — so the hard filter was reverted in favor of this soft guarantee instead.
✕ before
Query: "blackrock case studies." All 6 real BlackRock mentions in the corpus are classified Capabilities & Overview or Additional Materials — 0 survive into the Case Studies section.
✓ after
1 real BlackRock match guaranteed a slot in Case Studies — the single best-scoring one, not all 6, and only because nothing else naturally made it in.
15Sections ordered by real per-query relevance_format()NEW
Each section is already best-first internally, so its own top score is its real relevance to this query — sections sort by that descending instead of a fixed sequence.
in plain English
A section's top score is the highest score among the individual slides inside it. This step uses that one number per section purely to decide display order — nothing about the slides inside each section changes.
the logic here
Look at each section's own highest score.
Sort sections by that top score, highest first.
An empty section sorts last but is still included, not hidden.
example — real section order for "customer analytics" today
Data & Analytics top score 0.0978 ← leads today (4 slides)
Approach & Methodology top score 0.0833 (8 slides)
Capabilities & Overview top score 0.0527 (7 slides)
Case Studies top score 0.0154 (15 slides — 2 real + 13 verified)
16Coverage check against the WHOLE corpus_check_coverage() · _coverage_note()NEW
Checks the query's own distinctive words (excluding generic business terms) for at least one literal mention ANYWHERE in the 122,623-slide corpus — not just what got retrieved. Separates two failure shapes that look identical to a rep but need different honesty: a genuine coverage gap (nothing has it) vs. a ranking gap (the content exists, it just didn't win).
in plain English
If you search for something and the results look weak, this tells you WHY: either the corpus genuinely has nothing on that exact term (so the closest related material is the honest best available), or the content does exist somewhere but didn't rank into your results this time. A rep can't tell these apart from the results page alone without this check.
the logic here
Pull the query's own distinctive words, skipping generic business terms that would always find thousands of hits and say nothing useful.
For each one, check for at least one literal mention anywhere in the whole corpus (cached per engine instance).
If any word has zero real mentions anywhere, surface a plain note naming it — "nothing in the corpus uses that exact wording, showing the closest related material instead."
example — real trace, "customer analytics"
uncovered_terms = null — every distinctive word has real coverage somewhere in the corpus, no note shown. (Real example where this DOES fire: "Fan360 audience analytics" — zero mentions of "Fan360" anywhere in 122,623 slides, a genuine gap, not a ranking miss.)
17Response assembled_format()NEW
{query, intents, client, use_case, expanded_terms, groups: {...}, master_sheet: [], total_matches}. master_sheet is now ALWAYS empty — kept only for backward compatibility — since verified rows are already inline inside groups["Case Studies"] (step 13), each carrying its own verified: true/false flag instead of living in a separate list.
in plain English
This is the internal package of data the backend hands back — not something you'd see directly, just what powers the actual screen. groups holds the real cards, organized into the (now eight-category) sections from step 13, each carrying its own real relevance score and a verified badge flag where it applies. master_sheet is a legacy field, always empty now — real consumers should stop reading it.
the logic here
Take everything decided so far — the sections and their slides from steps 13-15, verified rows already merged in.
Package it into one response, each card carrying its own real title, a separate deck-name field, and a verified flag.
Hand it back to the chat assistant.
example — real assembled response (trimmed), "customer analytics" today
18Agent surfaces the resultspresent_groups (frontend tool)
The chat agent calls present_groups to render slide cards, in whatever order the backend already curated. The system prompt used to frame the team's master-sheet tracker as a "SEPARATE" source, nudging the agent to give it its own heading even post-merge — fixed to explicitly say verified rows belong in the SAME Case Studies section, mixed by relevance, with the per-card badge (step 13) as the only distinction.
in plain Englishpresent_groups is the internal name of the action the assistant takes to display organized, sectioned slide cards instead of one flat list. Two places used to quietly re-split Case Studies and Verified Case Studies back into separate headings even after the backend merged them — this step's own fallback grouping, and the wording the assistant itself was given — both fixed, confirmed live: a real query now returns one merged "Case Studies" section, zero as a separate "Verified" one.
the logic here
The assistant receives the assembled package from step 17.
It calls present_groups, handing over the sections in the exact order the backend already decided in step 15 — instructed not to re-sort them itself, and now also instructed not to re-split the merged Case Studies section back apart.
The chat window renders each section as its own labeled group of cards, in that order.
Real bug found and fixed here (2026-09-27): a second, completely separate code path — a hardcoded fallback heading for anything the assistant never explicitly placed anywhere — never filtered by content type at all, so demo cards and real slides the assistant left unmentioned rendered together, unsplit, under one literal "Other" label. Two other theories (a backend classification gap, a stale deploy) were tested and ruled out with direct evidence first. See version history below for the full trace. Status: pushed, merge/deploy pending as of this writing.
example — real production case, Goldman Sachs India CFO pitch
47 candidates retrieved, 16 explicitly named across 4 sections (incl. "Demos": 4)
31 left over, never mentioned by the assistant — 24 demo cards + 7 real slides, previously unsplit
after the fix:
Demos section = 4 + 24 recovered = 28
fallback heading = 7 genuinely uncategorized slides only
47 = 47 — every candidate accounted for, both before and after
A person can click any card to ask why it matched. The backend explain endpoint already existed, but was never called from the UI until this quarter — one batched Gemini call, given the same query this result set was actually ranked against (pulled from the message's own tool-call arguments, not the raw chat text, since the agent may have decomposed a multi-part brief differently).
in plain English
A Gemini call here just means asking Google's AI model to write the explanation text. "Batched" means all the currently-visible cards' explanations can be requested together in one single request rather than one at a time, which is what keeps clicking around fast rather than triggering a slow round-trip per card.
the logic here
Wait until a person actually clicks "Why this?" on a specific card — nothing happens automatically before that.
When clicked, look up the exact search text that particular result set was actually ranked against (from the assistant's own internal record, not necessarily your literal typed words).
Send that search text plus the card's own content to the AI model, along with honest instructions on how to describe a genuine match, a loose match, or no match at all.
Show the returned sentence or two directly on the card.
Fixed twice after shipping: first, cards from the master case-study sheet failed 100% of the time (no database row to explain); then, explanations were found to fabricate connections that don't exist on the slide at all.
Real bug found this week — the array can silently drop an entry mid-batch: the batched call asks the model to return exactly N explanations for N cards, then matches them back up by POSITION with no length check. Confirmed directly on a real 30-item batch: the model's array had one fewer entry than expected, so from item 13 onward every explanation silently described the WRONG (next) card instead — a full cascade, not just a truncated tail. This is a live production bug, not just an eval-measurement artifact — a rep could see a confident, wrong "why this" in front of a client. Fixed: if the returned array length doesn't match the item count, fall back to safe per-item explanations for the WHOLE batch, since positional trust is broken for all of it, not only the visible seam.
✕ before
Real 30-card batch. Model returns 29 explanations. From card 13 onward, every "why this" text confidently describes the NEXT card, not the one it's attached to.
✓ after
Array length is checked before trusting any position. 29 ≠ 30 → the whole batch falls back to a safe, generic per-card explanation instead of a confidently wrong one.
Real bug found this week — the explanation was reading a shorter snippet than the real slide text: the prompt was fed the same 200-character short_description used for the small UI card preview, while the real indexed snippet is 600 characters. Fixed by fetching the full 600-character snippet directly from the database for explanation/judging — the 200-character UI card display itself is unchanged, since that's a legitimate, separate display concern.
✕ before
Explanation generated from the first 200 characters of the slide's text — the same short preview the small UI card uses.
✓ after
Explanation generated from the real, full 600-character indexed snippet. Card preview length on screen is unchanged — this only affects what the AI reads to explain itself.
ENDPerson sees grouped, ranked, deduplicated slide cards
Content-type sections, ordered by relevance to the actual question asked, each internally ranked, each genuinely distinct — no boilerplate flooding a section, no deck flooding the response, no unrelated slide masquerading as "similar" because it happened to share a deck name.
the logic here
Nothing new happens at this last step — it's simply the sum of every decision made in steps 01 through 18b above, all now visible on screen at once.
What made it here survived every check along the way: it was in the search-scope pool before ranking even started (06), matched your actual wording or a genuinely related term (07a/07b), scored well on folder/type/recency (08), wasn't a duplicate (09/10), didn't crowd out other decks (11), cleared the relevance floor (12), landed in one of eight real content-type sections — with verified case studies merged in, never displacing a real match (13) — under both the section and whole-response caps (14), shown in the order that's genuinely most relevant to your question (15).
running example, end to end — real trace, 2026-09-22
"customer analytics" → 34 total matches across 4 sections, led by Data & Analytics, top card "Identified target customer groups for rolling out efficient Cross-sell/Up-sell marketing initiatives" at score 0.0978 (a genuine dashboard artifact — NPS distribution, revenue per account, customer journey — not just a slide that mentions data) — Case Studies merged 2 real slides with 13 verified rows from the master sheet into one section, none displacing the other.
Does it actually work? Real measured results
From the self-improving eval loop — real sales-team queries pulled from live usage logs, run through the actual pipeline above, judged correct / wrong / 50-50 by an independent AI grader. Goal is not 100% correct — that would mean the system got too conservative and started hiding genuinely good results. Goal is a real, evidenced reduction in "wrong," without ever creating a zero-result query.
Where the "wrong" verdicts come from, by section
Case Studies
388 (52%)
Capabilities & Overview
141
Approach & Methodology
126
Data & Analytics
53
Team & Credentials
19
Additional Materials
15
Credibility & Recognition
3
745 real wrong verdicts, one measurement pass. Case Studies alone accounts for over half — which is exactly why steps 13, 14b, and the claim-requirement gate all concentrate their real fixes on that one section specifically, not spread evenly across all eight. The same pattern — Case Studies as the dominant source of "wrong" — reproduced across every later measurement pass this quarter too.
Correct vs. wrong, session start vs. now
Before this quarter's fixes · 1,948 results judged
34%
42%
24%
Now · 3,116 results judged
47%
31%
22%
correct wrong 50-50 / partial
Real, evidenced movement — driven mostly by fixing genuine bugs (the negation carve-out, the explanation array-cascade, the evidence-truncation fix, the entity/phrase/person-name work above), not by loosening the grading. Zero-result queries stayed at zero throughout. Honest caveat: the "before" number is from an earlier 80-query set and the "now" number from a separately-cleaned 133-query set built after fixing a query-provenance bug in the measurement harness itself — not a perfectly matched pair, but each individual jump was independently verified against its own query set at the time.
SlideSearch — self-improving eval loop
Every run, what we found
The standing methodology, from a recorded call: pull real sales-team queries straight from live Supabase usage logs — not made-up ones — run each through the real pipeline shown in the other tab, fire the real "Why this?" explanation for every result, and have an independent AI judge grade each result+explanation as correct / wrong / 50-50. The explicit goal is not 100% correct — that would mean the system got too conservative and started hiding genuinely good results (the standing analogy: a bank driving loan defaults to 0% by refusing to lend isn't a success). Standing rule: a zero-result query is the worst possible outcome, worse than a tangential one.
Every run, in order — real numbers, nothing smoothed over
0 · Baseline reference point
_eval_loop_run1_PRE_ENTITY_GATE.json · 80-query set · 1,948 judged
34%
42%
24%
1 · Codex's hard exclusion filter reverted
_eval_loop_run1_CODEX_HARD_FILTER.json · 80-query set · 1,641 judged
36%
39%
25%
3 real queries returned ZERO results. The percentages above look almost fine in isolation — that's exactly why they're dangerous to read alone. Reverted immediately per the standing rule: zero results is worse than any wrong result.
_eval_loop_run1_SOFTBOOST_ONLY.json · 80-query set · 1,957 judged
34%
42%
24%
+0.8pp correct vs. baseline, 0 zero-result queries — safe, but not a real improvement. Called out directly: "I am not seeing improvement, how can you commit or push."
_eval_loop_run1_ENTITY_GUARANTEE.json · 80-query set · 1,978 judged
34%
41%
24%
Fixed the real "blackrock case studies" anchor case directly — verified working — but the whole-set aggregate barely moved again. This is the point the investigation stopped tuning entities and started asking whether the MEASUREMENT itself was trustworthy.
Pivot: three measurement-integrity bugs found, not retrieval bugs
A second independent AI reviewer ("Astra") was brought in deliberately at this point. Found three real bugs in how the eval loop itself measured quality — not in search ranking: (1) the harness was reading the wrong query file, including follow-on queries that skew the set; (2) the judge and the "why this" explanation were both being fed a 200-character truncated snippet, while the real indexed text is 600 characters; (3) the judge's own rubric produced an inconsistent verdict on compound-entity queries. All three verified directly against real code and real data before being trusted, then fixed.
+3.1pp correct in one pass — bigger than all three rounds of entity tuning above, combined. Came from fixing two real bugs (the "without case studies" carve-out, and the explanation array-cascade), not from any new ranking heuristic. This is the finding that reshaped the rest of the work: fix what's broken about how you're measuring before tuning what you're measuring.
Switch to a properly-built query set
The query-provenance bug above meant the 80-query set itself wasn't clean. A separately deduplicated, first-session-only, 133-query "standalone" set was built directly from real Supabase logs and wired into the harness. Every run below uses that set — not the 80-query one above, so treat the jump between eras as a set change too, not purely a quality change.
5 · Partial run, killed early for cost partial data
_eval_loop_standalone_PARTIAL73_PRE_PERSONNAME.json · 73 of 133 queries · 39 errors excluded
42%
32%
26%
Killed mid-run on direct instruction ("stop it, work with the 58") once cost became a live concern — worked with the real partial data already collected rather than burning more budget to finish it. Measurement fixes only at this point, no person-name detector or batched judge yet.
The full, clean 133-query run. Batching cut judge API calls roughly 10-15x with zero errors across the whole run — pure efficiency, verified not to cost rubric fidelity. Person-name detection verified working directly on real cases ("Anand S," partially on "Naveen Gattu" — an honestly-reported, narrower limitation, not hidden).
Real, individually-verified anchor-case win ("capital call reporting" now surfaces the right slide #1) — but essentially no aggregate movement. This is the third straight narrow term-detection mechanism (entity → person-name → phrase) to show this exact shape: a genuine local win, no whole-set movement.
correct wrong 50-50 / partial
what all seven runs, together, actually say1. Measurement bugs moved the needle more than retrieval tuning did. Three rounds of real, individually-verified entity/ranking work (runs 2, 3, and later run 7) each produced a flat or near-flat aggregate. One pass fixing how the eval loop itself measured quality (run 4) produced a bigger jump than all of that tuning combined.
2. A hard filter is tempting and dangerous. Codex's version (run 1) looked almost fine on the topline percentages and was still a real regression — 3 zero-result queries, the one outcome the standing rule treats as worse than anything else. Percentages alone don't catch that; you have to check for zero-result queries specifically, every time.
3. Diminishing returns from stacking narrow detectors. Entity → person-name → phrase, in that order — each one has a real, directly-verified anchor case, and each one left the whole-set aggregate basically where it found it. That's an honest signal that the next real gain probably needs something structural — a genuine query-aware reranker — rather than a fourth narrow pattern.
you'll get asked this
So is it actually better now, or did you just get better at measuring it?
Both, honestly, and the measurement part came first. Correct went from 33.6% to 47.1%, wrong from 42.4% to 30.6% — real, and zero-result queries stayed at zero the entire time. But the single biggest jump (run 4) came from fixing bugs in how "wrong" was being measured — a truncated snippet, a mismatched query file, an inconsistent judge rubric — not from a smarter ranking algorithm. That's not a knock on the result: catching that the ruler itself was miscalibrated, before trusting what it says about the thing being measured, is exactly the right order of operations. The three rounds of actual retrieval tuning after that were real too, just smaller in effect, and the last one (phrase detection) is flat enough that it's an open question whether a fourth narrow detector is worth building versus something more structural.