A collection of particularly difficult test scenarios for evaluating browser-use.
-
Updated
May 15, 2026 - HTML
A collection of particularly difficult test scenarios for evaluating browser-use.
A Claude Code skill that adds a rubric-based eval layer to any agent project. Framework-agnostic — generates rubric, test cases, judge prompt, and harness. Returns a weighted score plus a judge-leniency signal.
Unit tests for AI agents that browse the web.
Open-source Agent Skill for Claude Code that makes Claude reason with explicit probabilities, base rates, and falsifiers. Distilled from 51 books; first in 18/18 blind passes vs baseline Claude and GPT-5.5, and benchmarked against 14 rival skills and prompt techniques with full cost telemetry (round 5).
Open-source, self-hosted agent platform for banks and fintechs, declarative agents, banking guardrails, CI-gated evals, tamper-evident audit. Beta.
Proof Gradient is the agent evolution protocol where every run leaves proof, every proof selects intelligence, and every selected artifact evolves the network.
A governance layer for autonomous AI agents, proven on a life: 30+ scheduled agents made auditable, self-healing, and human-gated by single-writer field ownership. Architecture and patterns, no personal data.
Consultant-grade skills, workflow, and templates for designing, evaluating, launching, and iterating AI Native products.
The Agent Harness Path — a self-contained course on building, evaluating, and governing LLM agents (14 HTML lessons, 12 stdlib-only toy notebooks).
Agent skill for AI Systems — build model-powered behavior: agents, retrieval, evals, and guardrails
Claude Code plugin that encodes a metal-roofing manufacturer's domain knowledge as 7 eval-scored skills — pricing, SKUs, material estimation, colors, production, suppliers, inventory. Reusable template for building domain-expert plugins.
Prove your Agent Skill works before you publish it.
An autonomous LLM fund that scores its own stock predictions — Brier score, calibration curves, eval + ablation harness — and serves its full decision history over MCP. LangGraph pipeline with deterministic risk guardrails.
AI mock-interviewer that grades itself - adaptive technical screens for AI/ML/DS roles with an LLM-as-judge eval pipeline (candidate report + interviewer self-audit)
Production-oriented template for building AI agent skills as verifiable software components — offline eval harness, source-grounding validators, structured logs/traces/metrics, replay artifacts, and a CI quality gate. Runs fully offline with a deterministic mock model.
Static HTML viewer for agent messages, tool calls, and grader results.
Evals & performance for LLM. Workshop on how to test Agents before they break production
To associate your repository with the evals topic, visit your repo's landing page and select "manage topics."