jiabins0303 / birefnet-lite-1024-webgpu

huggingface.co
Total runs: 1.9K
24-hour runs: 4
7-day runs: 119
30-day runs: 1.6K
Model's Last Updated: October 02 2026
image-segmentation

Introduction of birefnet-lite-1024-webgpu

Model Details of birefnet-lite-1024-webgpu

BiRefNet-lite 1024 (ONNX, WebGPU-runnable)

A structurally patched ONNX export of BiRefNet-lite at 1024x1024 that runs on onnxruntime-web's WebGPU execution provider . The upstream export does not: it falls back to the CPU EP on most of its graph and dies of std::bad_alloc .

The weights are unchanged. Only the graph topology was rewritten, and both rewrites are proved bit-identical to the original on the CPU EP ( max|diff| = 0.000e+00 over the test set).

Credits and licence
  • Original model: ZhengPeng7/BiRefNet (MIT) — Peng Zheng et al.
  • ONNX export this was derived from: onnx-community/BiRefNet_lite-ONNX
  • Graph surgery: this repo. MIT, same as upstream.

If you use this, cite the BiRefNet authors, not us — the model is theirs.

Why the upstream export cannot run on WebGPU

Two independent blockers, which have to be fixed in this order because the first one masks the second:

  1. 59 Split nodes with 32 outputs each. WebGPU's maxStorageBuffersPerShaderStage is 8 by spec (10 on many adapters), so these never compile. Rewritten into trees of <=6 outputs.
  2. 80 GatherND nodes. deform_conv2d has no ONNX operator, so the exporter emulates it with GatherND — for which ORT's WebGPU EP has no kernel . All 80 ran on the CPU EP, which meant materialising a [1,1,64,49,256,256] fp16 im2col buffer — 392MB , several live at once — on ORT Web's 32-bit wasm heap . That is the std::bad_alloc ; the resolution itself was never the problem. Rewritten to Gather plus int32 index arithmetic, and variadic Sum to binary Add .

Host<->device copies drop from 100 to 80 and the graph runs end to end on WebGPU.

Two things that are easy to get wrong if you redo this:

  • Cast to int32 LATE. Doing the index arithmetic in int32 doubles the CPU<->GPU crossings (180 copies / 9.1s vs 80 / 3.4s). Do the maths in int64 and cast once, immediately before the Gather .
  • Use a sequential Add chain, not a balanced tree. fp16 addition is not associative; a balanced tree changes the summation order and breaks bit-identity with the original graph.
Measurements

Against BiRefNet-lite-512 on the same 30-image set, same compositor, same metric ( edgeSharpness().gradient ):

all-30 gradient products interiorSoft WebGPU time
BiRefNet-lite-512 0.0989 0.1095 0.00296 0.85s
this (1024) 0.1436 0.1754 0.00043 3.4s

+45% edge gradient for ~4x the time. It also recovers subjects 512 misses entirely — one test image returns an empty matte at 512 and 18.4% foreground here.

Usage (transformers.js)
const model = await AutoModel.from_pretrained('jiabins0303/birefnet-lite-1024-webgpu', {
  device: 'webgpu',
  dtype: 'fp32',            // the file is ALREADY fp16; this stops the library
  model_file_name: 'model_fp16',  // appending a dtype suffix to the filename
});

Requires an adapter reporting maxStorageBuffersPerShaderStage >= 8 . The graph emits logits , not an image — apply sigmoid, do not read it as 0-255.

Provenance

onnx/model_fp16.onnx sha256 4059896039dfccb0f15b9080ff06d11d90e499449bb045e797055eb8901cf5f4

Reproduce with patch_split.py then patch_deform.py , and verify with verify_patch.py (correlation gate plus max-abs-diff against the unpatched graph on the CPU EP).

Runs of jiabins0303 birefnet-lite-1024-webgpu on huggingface.co

1.9K
Total runs
4
24-hour runs
57
3-day runs
119
7-day runs
1.6K
30-day runs

More Information About birefnet-lite-1024-webgpu huggingface.co Model

More birefnet-lite-1024-webgpu license Visit here:

https://choosealicense.com/licenses/mit

birefnet-lite-1024-webgpu huggingface.co

birefnet-lite-1024-webgpu huggingface.co is an AI model on huggingface.co that provides birefnet-lite-1024-webgpu's model effect (), which can be used instantly with this jiabins0303 birefnet-lite-1024-webgpu model. huggingface.co supports a free trial of the birefnet-lite-1024-webgpu model, and also provides paid use of the birefnet-lite-1024-webgpu. Support call birefnet-lite-1024-webgpu model through api, including Node.js, Python, http.

birefnet-lite-1024-webgpu huggingface.co Url

https://huggingface.co/jiabins0303/birefnet-lite-1024-webgpu

jiabins0303 birefnet-lite-1024-webgpu online free

birefnet-lite-1024-webgpu huggingface.co is an online trial and call api platform, which integrates birefnet-lite-1024-webgpu's modeling effects, including api services, and provides a free online trial of birefnet-lite-1024-webgpu, you can try birefnet-lite-1024-webgpu online for free by clicking the link below.

jiabins0303 birefnet-lite-1024-webgpu online free url in huggingface.co:

https://huggingface.co/jiabins0303/birefnet-lite-1024-webgpu

birefnet-lite-1024-webgpu install

birefnet-lite-1024-webgpu is an open source model from GitHub that offers a free installation service, and any user can find birefnet-lite-1024-webgpu on GitHub to install. At the same time, huggingface.co provides the effect of birefnet-lite-1024-webgpu install, users can directly use birefnet-lite-1024-webgpu installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

birefnet-lite-1024-webgpu install url in huggingface.co:

https://huggingface.co/jiabins0303/birefnet-lite-1024-webgpu

Url of birefnet-lite-1024-webgpu

birefnet-lite-1024-webgpu huggingface.co Url

Provider of birefnet-lite-1024-webgpu huggingface.co

jiabins0303
ORGANIZATIONS