kernels-community / metal-flash-sdpa

huggingface.co
Total runs: 637
24-hour runs: 0
7-day runs: 38
30-day runs: 278
Model's Last Updated: May 05 2026

Introduction of metal-flash-sdpa

Model Details of metal-flash-sdpa

Metal Flash SDPA

Optimized SDPA kernels inspired by Flash Attention for Metal.

Supported Features
  • Variable-length sequences without padding
  • Causal masking
  • Grouped Query Attention (GQA) and Multi-Query Attention (MQA)
  • Softcapping support for attention score regularization
  • Data types: float32 , float16 , bfloat16
  • Head dimensions: 32 , 64 , 72 , 80 , 96 , 128 , 256
API Reference
flash_attention_varlen
metal_flash_sdpa.flash_attention_varlen(
    out: torch.Tensor,
    query: torch.Tensor,
    key: torch.Tensor,
    value: torch.Tensor,
    cu_seqlens_q: torch.Tensor,
    cu_seqlens_k: torch.Tensor,
    max_seqlen_q: int,
    max_seqlen_k: int,
    do_causal: bool,
    scale: float,
    softcapping: float
) -> None
  • out : Output tensor [total_q_tokens, num_heads, head_dim] , modified in-place.
  • query/key/value : Input tensors [total_tokens, num_heads(_kv), head_dim] .
  • cu_seqlens_q/cu_seqlens_k : Cumulative sequence lengths ( torch.int32 ), [batch_size + 1] .
  • max_seqlen_q/max_seqlen_k : Maximum sequence lengths.
  • do_causal : Enable causal masking.
  • scale : Attention score scaling factor (e.g., 1/sqrt(head_dim) ).
  • softcapping : Softcapping value for score regularization (use 1.0 for no softcapping).
flash_attn_varlen_func

Compatibility wrapper matching the original Flash Attention API:

out = metal_flash_sdpa.flash_attn_varlen_func(
    q: torch.Tensor,
    k: torch.Tensor,
    v: torch.Tensor,
    cu_seqlens_q: torch.Tensor,
    cu_seqlens_k: torch.Tensor,
    max_seqlen_q: int,
    max_seqlen_k: int,
    dropout_p: float = 0.0,
    softmax_scale: Optional[float] = None,
    causal: bool = False,
    window_size: Tuple[int, int] = (-1, -1),
    alibi_slopes: Optional[torch.Tensor] = None,
    deterministic: bool = False,
    return_attn_probs: bool = False
)

Runs of kernels-community metal-flash-sdpa on huggingface.co

637
Total runs
0
24-hour runs
0
3-day runs
38
7-day runs
278
30-day runs

More Information About metal-flash-sdpa huggingface.co Model

More metal-flash-sdpa license Visit here:

https://choosealicense.com/licenses/apache-2.0

metal-flash-sdpa huggingface.co

metal-flash-sdpa huggingface.co is an AI model on huggingface.co that provides metal-flash-sdpa's model effect (), which can be used instantly with this kernels-community metal-flash-sdpa model. huggingface.co supports a free trial of the metal-flash-sdpa model, and also provides paid use of the metal-flash-sdpa. Support call metal-flash-sdpa model through api, including Node.js, Python, http.

kernels-community metal-flash-sdpa online free

metal-flash-sdpa huggingface.co is an online trial and call api platform, which integrates metal-flash-sdpa's modeling effects, including api services, and provides a free online trial of metal-flash-sdpa, you can try metal-flash-sdpa online for free by clicking the link below.

kernels-community metal-flash-sdpa online free url in huggingface.co:

https://huggingface.co/kernels-community/metal-flash-sdpa

metal-flash-sdpa install

metal-flash-sdpa is an open source model from GitHub that offers a free installation service, and any user can find metal-flash-sdpa on GitHub to install. At the same time, huggingface.co provides the effect of metal-flash-sdpa install, users can directly use metal-flash-sdpa installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

metal-flash-sdpa install url in huggingface.co:

https://huggingface.co/kernels-community/metal-flash-sdpa

Url of metal-flash-sdpa

Provider of metal-flash-sdpa huggingface.co

kernels-community
ORGANIZATIONS

Other API from kernels-community