T
traeai
登录

概念

Terminal-Bench

别名:TB

评估模型终端任务性能的基准测试套件

已跟踪 7 条高相关材料

TraeAI 观察

相关材料

已收录 7 条与 Terminal-Bench 相关的内容,按评分排序。

Latent Space 图标

[AINews] not much happened today

Latent Space1843 字 (约 8 分钟)
85

DeepSeek V4-Flash 0731通过微调实现性能跃升,成本降低60%,但未改变架构。

入选理由:DeepSeek V4-Flash 0731在Terminal-Bench测试中提升25.8个百分点至82.7

精选文章#DeepSeek#AI模型#API#技术更新中英混合
Vercel News 图标

DeepSeek V4 Flash now runs updated weights on AI Gateway

Vercel News234 字 (约 1 分钟)
85

DeepSeek V4 Flash在AI Gateway更新权重后Terminal-Bench得分提升至82.7,增强代理能力且无需代码修改。

入选理由:Terminal-Bench测试得分从56.9提升至82.7,性能提升25.8分

精选文章#AI模型#AI Gateway#DeepSeek V4 Flash#终端测试英文
OpenAI Just Introduced GPT 5.6 (Beats Claude Fable 5 And Mythos)

OpenAI Just Introduced GPT 5.6 (Beats Claude Fable 5 And Mythos)

TheAIGRID5259 字 (约 22 分钟)
85

OpenAI 推出 GPT 5.6 系列模型,包含 Soul、Terror 和 Luna,Soul 超越 Claude Mythos 5 在终端任务表现。

入选理由:GPT 5.6 Soul 在终端任务中超越 Claude Mythos 5。

精选视频#GPT#AI模型#OpenAI#Claude英文
1/ Today at #GoogleIO, we’re releasing Gemini 3.5, our latest family of models combining frontier in...

Jeff Dean 发布 Gemini 3.5

Jeff Dean(@JeffDean)268 字 (约 2 分钟)
85

Google 发布 Gemini 3.5 模型家族,首发 3.5 Flash 专注于复杂智能体工作流,在编码和代理基准测试中超越 3.1 Pro,速度比前沿模型快 4 倍,在 Antigravity 中优化后可达 12 倍。

入选理由:Gemini 3.5 Flash 专为执行复杂、长周期的智能体工作流而设计。

精选推文#Google#Gemini#AI Agents#LLM#Google I/O英文
I'm very excited about this extension to the celebrated Terminal-Bench to science.

If you're a scie...

Thomas Wolf is excited about the extension of Terminal-Bench to scientific fields, known as Terminal-Bench Science. This benchmark evaluates AI models' ability to control tools via the command line to achieve scientific goals. It's open for contributions of real scientific workflows until August 2026, aiming to improve AI models' assistance in research work.

入选理由:Terminal-Bench Science evaluates AI models' performance in handling scientific workflows through command-line tools.

精选推文#AI#Science#Terminal-Bench#Benchmarking#Command Line英文
What's the tea on harnesses?

什么是 Harness?

LangChain269 字 (约 2 分钟)
72

Harness 是构建 AI Agent 的核心基础设施,由工具、执行环境、系统提示词和文件系统组成。通过优化 Harness 工程(如调整上下文和提示词),开发者可以在不更换底层模型的情况下显著提升 Agent 在特定基准测试(如 Terminal Bench)中的性能。

入选理由:Harness 定义为模型访问的工具、执行环境、系统提示词和文件系统的集合。

精选视频#AI Agents#Harness Engineering#LLM#LangChain英文

跨材料问答 · Terminal-Bench

回答基于:Terminal-Bench 相关 7 条材料
    0 / 500

    AI 可能会生成不准确的信息,请核实重要内容