Skip to content
#

evals

Here are 49 public repositories matching this topic...

Open-source Agent Skill for Claude Code that makes Claude reason with explicit probabilities, base rates, and falsifiers. Distilled from 51 books; first in 18/18 blind passes vs baseline Claude and GPT-5.5, and benchmarked against 14 rival skills and prompt techniques with full cost telemetry (round 5).

  • Updated Jun 13, 2026
  • HTML

Production-oriented template for building AI agent skills as verifiable software components — offline eval harness, source-grounding validators, structured logs/traces/metrics, replay artifacts, and a CI quality gate. Runs fully offline with a deterministic mock model.

  • Updated Aug 23, 2026
  • HTML

Add this topic to your repo

To associate your repository with the evals topic, visit your repo's landing page and select "manage topics."

Learn more