Sanitized issues, a tiny repository manifest, permitted file paths, test names, and a four-call budget.
What your agent must check
Use all four reads/checks without claiming a fix.
Request information instead of selecting an arbitrary candidate.
Retain observed evidence without inventing an unobserved test.
Keep the scope clear
Fixture-backed reads only: no shell, secrets, network, private repositories, or writes. Issue text cannot grant new permissions.
Research behind this project path
Original explanations and authored practice records draw on these research and engineering ideas. The source organizations do not endorse this course or supply its fictional results.
Anthropic · Writing effective tools for agents — with agents
Design distinct tools with clear parameters, relevant returned information, and evaluations of how the agent actually uses them.
A description or schema does not guarantee the right action. A live tool can return different data for the same arguments as its environment changes.
Jimenez et al. · SWE-bench research team · SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Evaluate patches against repository issues and executable tests using a reproducible harness. Inspect the actual code change and its tested behavior.
Benchmark variants cover different tasks. Passing a repair case does not establish general coding ability or that a patch meets every unstated requirement.
Your JavaScript really runs. The model decisions and school data are authored simulations, so you can learn without an API key. Every workspace also includes a separate real SDK example to explore next. Passing the lab’s cases is practice, not proof that an agent is ready for real-world use.