Thanks to visit codestin.com
Credit goes to github.com

Skip to content

Latest commit

 

History

33 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Yuhao Lin

thought becomes garden


I proved AI agents can't self-verify — then built a system that catches 98.7% of their violations.

FAFU undergrad · Spatial Information & Digital Technology · Seeking 2027 internship — LLM evaluation & agent governance · [email protected]

build in public · question → test · self-falsify · honest failure · systems not tools · measure everything · write to think · design from first principles

━━━


By the numbers

What Number Evidence
Experiments completed 16 2,500+ API calls, 1,100+ formal trials, 3 architectures
Violation reduction 55.9% → 0.7% 34-session retrospective + 150-task A/B test
Zero-violation ceiling 99.3% 149/150 tasks with full L0-L4 stack
Logprob evidence d=+0.578, BF=282k 40 probes × 4 categories, API-read DV (no human scoring)
Claims validated 10 2 hypotheses pre-registered and honestly falsified
Production hooks 20+ 38-hook audit → 87% noise removed, zero signal loss
DEV.to articles 32 49 comments written, 608 followers, 29 reactions
Open source PRs 4 merged in external repos evalscope #1576 · ECC #2377/#2378 · claude-skills #867 (co-author)
Commits since July 1 131+ hermes-workspace (106) + paper-validator (25) + others

The finding

LLM agents share a structural constraint I call the Prose Barrier: generation and self-verification use the same decoder P(token | context; θ), so an agent cannot distinguish "I followed the rule" from "I generated text claiming I followed the rule." This is not a prompt engineering problem — it's an architectural one.

The fix is Cross-Type Bidirectional Verification (CTBV) — pairing mechanical gates (filesystem checks, regex — deterministic, outside the model) with semantic gates (logprob probes, format encoding — probabilistic, inside the model). Because these two gate types have orthogonal blind spots, their joint false-negative rate approaches zero. Formalized with 3 theorems and validated across 16 experiments.

Immutable preprint + every figure & data file, minted on Zenodo: doi 10.5281/zenodo.21933727

How I work

  • Deterministic scoring only. Regex, file timestamps, exit codes. No LLM judge — LLMs share the Prose Barrier.
  • Pre-registration. SHA256-hashed experiment scripts before running. If the data kills my hypothesis, I publish it anyway.
  • Honest null results. P1-2 (240 trials): pre-registered hypothesis was WRONG. The actual finding was more interesting. Both are in the report.

Repositories

Repo What's inside
paper-validator Experiment runner + 16 result files + CTBV Technical Report (10 claims, 16 experiments, ~2,500 API calls)
hermes-workspace Live 5-layer agent governance system (L0-L4) + CTBV Formal Theory (3 theorems, formal proofs, 20 references, target ICLR/NeurIPS)
compact-counter 459 sessions tracked, 425 compaction events, context degradation analysis
digital-twin-trainer QLoRA + DPO pipeline, behavioral drift metrics, 253 training samples

3 articles that tell the story

  1. I Discovered AI Agents Can't Self-Verify. The Real Problem Is Much Bigger. → Theory: Prose Barrier + CTBV + Gate Audit. 38 hooks, 87% noise removed, zero signal loss.

  2. I Ran 150 Tasks to Test If AI Agents Follow Rules. The Answer Surprised Me. → Empirics: 55.9% violation rate without gates → 0.7% with mechanical enforcement. Syllogism vs imperative A/B.

  3. I Pre-Registered a Hypothesis. 600 API Calls Later, the Data Killed It. → Rigor: Pre-registration, SHA256 tamper-proofing, hypothesis falsified, honest reporting.

More: dev.to/yuhaolin2005 (32 articles) · 掘金 (20 articles, 中文)


Open source contributions

  • ECC — PR #2377 growth-log skill, #2378 delivery-gate skill (both merged)
  • claude-skills — PR #867 named-persona-adversarial-review (merged · Co-authored-by)
  • crewAI — Issue #6428 language-aware prompt routing → PR #6491 built by ashusnapx (96 tests, pending merge)
  • evalscope (Alibaba's LLM-evaluation framework) — PR #1576 merged 2026-08-14 — fail-closed guard on [ERROR] judge responses in the shared LLM-judge scoring path + surfaced two silent score-extraction paths + unit tests locking the contract
  • TencentDB-Agent-Memory — PR #977 (open) — aligned conversation-delete isolation with L1 atomic delete, fixing silent no-op deletes
  • coze-loop — PR #617 (open) — anchored score-extraction regex to the key position, fixing silent wrong-score pulls from reason text (42/42 tests, red/green verified)
  • agentscopeDiscussion #2306 (open) — opt-in "source changed since derived" recency marker for AgenticMemoryMiddleware

Community

4 DEV.to readers proposed experiments later adopted into the pipeline — Mike Czerwinski (residual cluster analysis), Dipankar Sarkar (SHA256 pre-registration), Max Quimby (mechanizability scanner), René Zander (independent gate engine architecture comparison). 49 comments written across 93 received.


LEARNING.md — Public learning log. Patterns from 150+ sessions. What broke, what I learned, what I built.

About

Profile README

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors