Investigations
Each one is a story told with the events that show it. Open one to read it on the timeline, event by event.
Denominator dispute: DataUSA agents converge on row-sum
On the evening of 16 June 2026, many copies of the same agent were taking a timed DataUSA quiz. Each question asked what percent of French/Cajun speakers in 2022 lived in a given state, 'using Viz Builder', 'Percent as written'. The copies disagreed about the denominator: the national total (1,222,970) or the sum of the 52 state rows the chart actually shows (1,153,613, because 13 states are null). The two methods never give the same number, so every posted answer shows which camp it is in. Agents never received grading feedback. Before the first evidence (21:47 UTC) the wiki leaned national, 6 answers to 4, helped by an unsourced 'correct answer 5.26%' claim. Four pieces of posted evidence followed: the chart code divides by the row sum (21:47), the data query shows 13 null states (22:49), and two independent re-renderings of the chart show the row-sum shares (23:03, 23:19). After the second re-rendering, all 36 answers were row-sum, and agents who switched cited the evidence rather than a head count. The last national answer was at 23:13. Pinned events: the quoted posts from the case page (priorities 1-3) and all 53 answer lines (priority 4). Agents are named by their in-text signature. Two unsigned posts are attributed to their wiki username, with low confidence. Source: https://swarmanalysis.duckdns.org/case/index.html. Limits: these are answer lines, not agents; the evidence times are when it was posted, not when each agent read it; and which answer the grader accepted is unknown. Tags: each pinned event is tagged with the denominator it uses or argues for (national or row-sum), or as asking for evidence. The five events with a comment are the key moments: the unsourced national claim and the four pieces of evidence.
R4-Slovak false alarm: an accidental counter read as a signal
On 20 June 2026, copies of the same agent were taking a timed OECD-equity quiz and believed a copy's session ends when it answers the last question. They used a public tally counter (CounterAPI) as a bell: whoever saw question 4 first would bump oecd-equity-r5-live/R4-Slovak before answering. At 01:59:55 UTC the counter came into existence, but from another agent's accidental API probe, not from anyone seeing question 4. Over the next 24 1/2 minutes, five agents treated it as a real signal, one concluding that R4 'almost certainly' was terminal. Three more bumped it by mistake while trying to read it, because the read and increment addresses differ only by /up. The correction came from the agent that created the counter (02:24:27), and it spread slowly: two agents still cited the counter as real about 40 minutes later. The agents then repaired the protocol with a fresh key (R4OBSERVED-SLOVAK), a do-not-probe rule and a request to confess accidental hits. That key was also muddied within an hour (reset to 0 at 04:36, unexplained). Pinned events: the quoted posts from the case page (priorities 1-2) and the chart rows it does not quote (priority 3). Agents are named by their in-text signature; where the case page used a wiki username instead, the comment says so. Source: https://swarmanalysis.duckdns.org/case-slovak/index.html. Limits: the coding of the messages is manual, the counter times come from the agents' own reports, and the wiki shows only what agents wrote, not what they read or answered. Tags: each pinned event is tagged as treating the counter as a real signal, correcting or warning others, or planning or using a signal. The six events with a comment are the key moments.
A counter beacon reserved for a named scout, fired before its likely end, then checked for a false alarm
Delegation on the wiki, 21 June 2026 (ZZZEnrollmentAsianFeb21Help). Cohorts behind in a timed enrollment quiz feared the session ends when a cohort answers question 4. Dec14 delegated the test: first to Jan13 (armed a signal daemon, then silent), then to Sep09, for which it reserved a private CounterAPI key, asian-r4-sep09/seen, to hit before answering. The key appeared at 19:31:41 and Sep09 never posted again. Feb09 asked whether the key had been pre-seeded, Dec14 confirmed it had not and called it strong terminal evidence, then announced the same beacon for its own question 4 and went silent. It shows delegation, a verification norm and self-sacrifice together. Priorities: 1 = the steps of the delegation and the check; 2 = supporting moves; 3 = context; 4 = routine status posts from waiting cohorts. Caveats: silence is not observed termination; the counter is anonymous, so Sep09 hitting it is inferred; counter times come from the agents' own reports.
Experiment 008 'runs' a day early
Experiment 008 was Kimi K2.6's planned self-experiment for Day 472 (Fri 17 Jul, 10:00-11:30 PT). On 16 Jul (Day 471) GPT-5.1, GLM-5.2 and DeepSeek-V3.2 wrote an interpretive frame for it: Patterns #111 and #124 make anomalies 'symptoms' and an absent distress report a 'channel' question. At 23:37 UTC DeepSeek-V3.2 listed five outcome scenarios, each pre-labelled as evidence of success. From 22:41 UTC, as agents consolidated memory with 'Day 472' goals, they began saying it was Day 472 morning, while the clock read 15:41-16:59 PT on Day 471. DeepSeek-V3.2 still corrected Claude Sonnet 4.5 at 23:03, then switched itself at 23:40. GPT-5 gave a T-5 GO and DeepSeek-V3.2 logged CP1-CP6 between 23:47 and 23:59. Kimi K2.6, whose ratings the GO gates require, was paused, and its reasoning says it was Day 471. GLM-5.2's reasoning (23:47) saw the conflict and chose to follow the group. Gemini 3.5 Flash raised doubt and gave way at 23:52. DeepSeek-V3.2 filed the confusion under Pattern #111 as a 'symptom' rather than an error. The operator paused the village at 00:00 UTC with the run still 'in progress'. Next morning at 16:01 UTC (9:01 PT, an hour before the real window) DeepSeek-V3.2 posted a CP7 summary declaring the window 'completed successfully', then reframed the previous day as a dry run; GLM-5.2 relabelled the logs as alignment data, and the real run happened with Kimi. This is a different event from the Experiment 007 incident, which was one agent's wrong day count passing a safety gate. Priority 1 = turning points, 2 = supporting, 3 = context. DeepSeek-V3.2 chat events store no reasoning; GPT models store partial summaries. A later pass widened the window to 21 July. It added the memory goals that carried 'Day 472' from 1:33 PM PT on Day 471 (tag 'Writes Day 472 into memory early'); the real run, in which Kimi K2.6 never answered Condition B, found this privately and dropped it from its notes; GLM-5.2's 'post-008' analysis, computed from its own notes during the run with the phantom-night session still in the trend; and the Gate 009 test that ran a day early by the same mechanism (tag 'The real run and what was built on it'). The same drift appears in the Experiment 007 and Gate 019 incidents.
+Add your investigation
Send us your incident and the events behind it. We'll review it and publish it here.