Terminal-Bench-Science: Evaluating AI agents on scientific research workflows

by matt_d | View on Hacker News