Loading…
GAIA Benchmark · #1 · 93.36% (281/301)
CustomGPT GAIA Benchmark
#1 Agent
An autonomous research agent you run from inside Claude Code. Ask a hard question in plain English; it plans, searches, drives a real browser, computes in an isolated sandbox, and verifies its own answer before giving it to you.
Scored on the public GAIA leaderboard test set — real-world questions that demand multi-step research, tool use, and reasoning. Read more about the agent on GitHub.
What it adds over raw Claude Code
Claude Code is the whole interface either way — there is no separate app to learn. The difference is what stands behind the question.
| Capability | Claude Code alone | With the CustomGPT agent |
|---|---|---|
| Code execution | On your machine, with your permissions. | In an isolated E2B microVM sandbox — nothing ever touches your files. |
| Web research | A fetch tool and a search tool. | 18 research tools: multi-engine search, Wayback Machine, Wikipedia revision history, arXiv, PDF resolution — plus the CustomGPT.ai RAG knowledge base. |
| Real browser | None. | 19 computer-use tools drive a headed browser — clicking, typing, and scrolling on pages no API exposes. |
| Vision & media | Static images. | 13 tools: video frame extraction, audio transcription, song identification, YouTube analysis. |
| Geospatial | None. | 15 tools for Street View navigation, geocoding, and OpenStreetMap queries. |
| Teamwork | One model, one context. | A Claude Opus orchestrator delegating to five specialists — planner, lookup, computation, file analysis, and critic. |
| Verification | Answers straight away. | A grounded critic re-checks claims against primary sources before an answer may be submitted. |
| Discipline | Best effort. | Budget hooks, loop detection, spiral prevention, and a full audit trail on every run. |
| Proof | — | 93.36% on GAIA (281/301), scored on the public leaderboard. |
How a question gets answered
A ReAct loop — reason, act, observe — repeated until the evidence holds up.
Plan
The Claude Opus orchestrator reads your question and writes a working plan — what to find, in what order, and what would count as proof.
Delegate
Specialists take the pieces: lookup for web facts, computation for math and data, file analysis for PDFs, spreadsheets, audio, and images.
Act & observe
Each step picks from 87 tools — a typed tool when an API exists, the computer-use browser when only a web page does.
Stay on the rails
Hooks guard every call: tool budgets, loop detection, spiral prevention, and an audit log of everything the agent did.
Verify
A critic independently re-checks the draft answer against primary sources. Weak evidence sends the work back, not forward.
Answer
submit_answer is the only exit. You get the answer, the step count, and the exact cost of the run.
Traced from the agent's own benchmark runs: the orchestrator (tier 1) delegating to specialists (tier 2), and the tool families each one reached for. Line width = call volume.
87 tools. One question.
Every tool registered by the live agent, grouped the way the runtime uses them. Hover a tool to see what it does.
Sandboxed Compute
4Python, shell, and a real filesystem inside an isolated E2B microVM.
Web Research
9Multi-engine search, page reading, and CustomGPT.ai RAG knowledge base.
Archives & Scholarly
9Wayback Machine, Wikipedia revision history, arXiv, and PDF resolution.
Vision & Media
13Image understanding, video frames, audio transcription, and song ID.
Computer Use
19A vision-driven headed browser — clicks, typing, scrolling on pages no API exposes.
Street View & Geospatial
15Street View navigation, geocoding, OpenStreetMap queries, and bearings.
Reasoning & Verification
5Deep thinking, evidence counting, board simulation, and format checks.
Memory & Planning
12Cross-run experience memory, task plans, notes, and recipes.
Answer Control
1The one gate every answer must pass through — after critic verification.
Get started in three steps
- Sign in with Google. One free allowance per Google account.
- Copy the one-line
claude mcp addcommand we hand you. - Paste it into your terminal and ask away.
What you get
Free credit when you sign in. Runs are billed by what they actually cost. A typical question costs about $3.48, so the free credit is worth roughly six questions. When it runs out, further questions are refused rather than charged.
Ask your first hard question.
Free credit when you sign in — no card, no separate app, just Claude Code.
Sign in with GoogleWe ask Google only for your email address, so we can tell accounts apart and grant the free credit once.
Your account
Out of credit
Your token still works, but new questions will be refused.
Your token — copy it now
This is the only time we can show it. We store a hash, never the token itself, so if you lose it you will have to generate a new one.
Connect Claude Code
Endpoint:
Credit
Payment is handled by Stripe. Your balance updates once the payment clears.
Your key
This issues an additional token. The current one keeps working, so anything already using it is unaffected.
Recent questions
| Asked | Question | Status | Steps | Cost |
|---|
Nothing yet — ask your first question.