Live features
12
8 healthy · 3 drifting · 1 deprecated
Calls · 24h
42,184
+12% vs prior 24h
Drift detected
2
brand-fit · gps-rover
Avg cost / call
$0.08
−$0.02 vs last week
§ III Feature detail · memetic-resonance-scorer drift detected · −3.2% accuracy over 24h trailing

Accuracy vs baseline · 30-day trailing

The classifier sits below its baseline for the first sustained stretch since launch. The drop is concentrated in cold-start prompts: queries that arrive without conversational context. The system has auto-generated an investigation; the next step is yours.

BASELINE · 92% THRESHOLD · 87%

Investigation · auto-generated

Failing categories
  • · cold-start prompts (no prior context), 38% failure rate
  • · prompts referencing post-2025 memes, 22% failure rate
  • · prompts in informal register, 14% failure rate
Hypothesized cause

Training data ends Q1 2026; post-Q1 meme corpus has shifted register significantly. Brand-fit constraint inherited the older register's assumptions. Two prompt revisions and an evaluation-set expansion would likely close the gap.

§ IV Optimization backlog profile-guided, sorted by impact
HI
Expand evaluation set with post-Q1 2026 meme register
memetic-resonance-scorer · would lift cold-start accuracy by an estimated +5 to +8%
+5–8% acc · 2h effort
HI
Add cold-start constraint to brand-fit prompt
memetic-resonance-scorer · explicit handling when no prior context exists
+3% acc · 30m effort
MD
Cache RAPTOR summaries by KB to cut retrieval latency
/retrieve/v2 · would reduce p50 by ~18% on repeat queries
−18% p50 · 4h effort
MD
Promote claim-temporal-arc to platinum tier
12 weeks of stable performance · 3 improvement cycles complete
infra priority · 5m effort
LO
Reduce token output cap on contradiction-explainer
analyzed runs · 142 outputs were ≥80% redundant with the structured fields
−12% cost · 15m effort