The DeepSeek visual model's multi-modal Agent ability is close to Opus 4.8, 4 items are 2:2

source··17:36 编辑

In comparison, DeepSeek announced the first batch of V4-Flash-Vision-Exp agent scores. There were 2 wins and 2 losses against Opus 4.8 in 4 multi-modal Agent reviews: Aggregation Last Exam 27.3 versus 25.7, ZeroBench 35.0 versus 34.0; ApexBench 36.5 versus 39.4; and Chartography 64.3 versus 65.0.

Compared to the plain text version of V4-Flash-0731, the visual version increased from 26.2 to 36.5 in ApexBench and from 25.2 to 27.3 for Aggregated Last Exam. The official statement indicates that the plain text version will ignore multi-modal content, so it mainly reflects the ability to supplement visual input. Coupled with the fact that the ability of the post-visual text agent was not significantly reduced, 6 of the 7 text reviews were superior to V4-Flash-0731. For example, DeepSWE upgraded from 54.4 to 59.3 (58.0 over Opus 4.8), and Toolathlon 75.9 almost tied with 76.2 of Opus 4.8.

The results are from DeepSeek's official self-test, not a third-party independent list; the public Code Agent text task uses the DeepSeek Harness minimal mode, and the inference level is max.

Original Link
说明: All Bitpush articles reflect the author's views only and do not constitute investment advice.

Related

Loading...