Skip to main content
Fidian
Research Request Early Access Early Access

Research

  • Three robots reach different heights on a mechanical course built into an old terminal.

    August 26, 2026

    Who is at the frontier of terminal tasks?

    We ran the top 20 models from Terminal-Bench 2.1 on TB-fn, our harder variant of the same 89 tasks. Only OpenAI and Anthropic held the top tier, and every model cost more per task.

  • A dial on a floor reads 75.2% while, in the crawlspace below, a Fidian robot inspects two gears labelled PASS and FAIL by lamplight.

    June 15, 2026

    Did your agent really improve?

    We rebuilt Terminal Bench 2 into task variants that expose broken verifiers, measure real robustness, and guide concrete improvements to the agent harness.

© 2026 Fidian, Inc.
  • Research
  • Contact
  • Privacy Policy
  • Terms & Conditions