- GPT0%
- OPUS0%
- GLM0%
Dynamic Beating AI News Flash: Terminal-Bench has released version 4.0, recalibrating the time, CPU, and memory when the agent performs tasks, and fixing 19 tasks. It has removed 8 tasks with saturation, denial, openly disclosed solutions, or quality issues. The maximum execution time for all tasks has been standardized to 8 hours, mainly to reduce interference from timeouts and environmental issues on performance.
In the latest leaderboard, Opus 5 + Claude Code ranks first with 51.8%, followed by Fable 5 at 44.5%. GLM-5.3 + Claude Code scored 41.8%, claiming the third spot, surpassing GPT-5.6 Sol + Codex at 37.3%. Among the top three, GLM-5.3 is the only model not from Anthropic.
In Terminal-Bench 3.0, GLM-5.3 ranked fourth with 32.4%, trailing GPT-5.6 Sol at 34.6%; by version 4.0, GLM-5.3 has climbed to third place, leading Sol by 4.5 percentage points in turn.
면책 조항: 현재 콘텐츠는 제3자 관점에서 제공되거나 제3자 관점에서 AI가 직접 번역한 것입니다. CoinEx는 콘텐츠의 진위성, 정확성, 독창성을 보장하지 않으며 CoinEx의 투자 조언으로 간주하지 않습니다. 암호화폐 가격은 변동성이 크므로 잠재적인 위험에 유의하시기 바랍니다.
- 코인가격24시간 변동