False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Abstract
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.
Community
TL;DR: In self-evolving search agents (a proposer writes questions + pseudo-labels from source documents, a solver answers them, and their agreement is the training reward), the proposer and solver can converge on shared errors: internal reward keeps rising while real correctness stalls. We call this co-cheating, measure it with a source-grounded post-hoc audit, and fix it with CrossFit.
CrossFit splits the proposer's source documents into two folds and scores each fold's questions with an auxiliary solver trained only on the other fold, so a same-source pseudo-label can never be echoed back as reward. The main solver still trains on everything; only the feedback path changes.
Results (Dr. Zero loop, Qwen3.5-4B / 9B):
- Co-cheating grows over rounds of self-evolution: false-agreement mass reaches 6.1% / 8.8%.
- Multi-sample verification (MSV) only trims it to 5.7% / 7.2%, at 6 extra labeler generations per candidate.
- CrossFit cuts it to 3.0% / 3.7%; replaying identical proposals with source-excluded feedback drives it to 0.4% / 0.1%.
- Downstream: +8.8 / +8.4 average points over coupled self-evolution and +8.7 / +7.8 over Search-R1 across seven search QA benchmarks.
Takeaway: when a model grades itself, track where the grader's supervision came from, not just whether the labels look right.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Dr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search Agents (2026)
- SearchMaster: Grounded and Regulated Self-Play for Search Agents (2026)
- CoEvoKG: Co-Evolving Knowledge Graphs with Self-Evolving Search Agents (2026)
- Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning (2026)
- CoEvo: Oracle-Grounded Self-Evolution of a Single Model for Multi-Step Causal Reasoning (2026)
- UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning (2026)
- The Chase Is the Curriculum, the Capture Anchors the Credit: Pursuit-Evasion Self-Play for Zero-Data LLM Reasoning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.39102 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper