Every
*_proj.weight
in the 80 decoder layers is stored as
float8_e4m3fn
with a sibling
*_proj.weight_scale
(float32, one per output row,
amax / 448
); dequantize as
weight.to(bf16) * weight_scale[:, None]
. Embeddings, LM head and norms are unchanged bf16. There is no
quantization_config
; load the layers yourself (see the Space that uses it).
Prompt format and sampling are unchanged from the original: see its card.
Runs of multimodalart computer-10-fp8 on huggingface.co
335
Total runs
0
24-hour runs
7
3-day runs
7
7-day runs
7
30-day runs
More Information About computer-10-fp8 huggingface.co Model
computer-10-fp8 huggingface.co is an AI model on huggingface.co that provides computer-10-fp8's model effect (), which can be used instantly with this multimodalart computer-10-fp8 model. huggingface.co supports a free trial of the computer-10-fp8 model, and also provides paid use of the computer-10-fp8. Support call computer-10-fp8 model through api, including Node.js, Python, http.
computer-10-fp8 huggingface.co is an online trial and call api platform, which integrates computer-10-fp8's modeling effects, including api services, and provides a free online trial of computer-10-fp8, you can try computer-10-fp8 online for free by clicking the link below.
multimodalart computer-10-fp8 online free url in huggingface.co:
computer-10-fp8 is an open source model from GitHub that offers a free installation service, and any user can find computer-10-fp8 on GitHub to install. At the same time, huggingface.co provides the effect of computer-10-fp8 install, users can directly use computer-10-fp8 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.