Ant Ling-3.0-Flash is officially open source, FP8 version is only 128GB

source··15:36 编辑

Comparative news, according to monitoring, InclusionAI (InclusionAI) officially released Ling-3.0-flash weights, and also provided the original BF16 and FP8 quantitative versions. Both versions use the MIT license, have launched Hugging Face and ModelScope, and can be self-deployed using SGLang or vLLM. The BF16 copyright weighs about 255GB, the FP8 version is about 128GB, and the size is nearly halved. FP8 uses lower accuracy to save parameters, which can lower the storage and video memory threshold. Of the 4 officially listed tests, the maximum score difference between FP8 and BF16 was 1.57 points. LING-3.0-Flash has a total parameter of 124 billion, and only 5.1 billion are activated per generation. The model supports 256,000 token contexts and is mainly aimed at agent tasks such as programming, search, in-depth research, and tool calls. According to official reviews, it met or surpassed the trillion-parameter predecessor Ring-2.6-1T on most benchmarks.

Original Link
说明: All Bitpush articles reflect the author's views only and do not constitute investment advice.

Related

Loading...