Pulpie Orange Base
extracts the main content from raw HTML, stripping navigation, ads, sidebars, and footers. It is an encoder that labels every HTML block as content or boilerplate in a single forward pass, so it approaches state-of-the-art extraction quality while running far faster and cheaper than autoregressive extractors.
At 610M parameters it sits between
Orange Small
and
Orange Large
, scoring 0.863 ROUGE-5 F1. For most use cases Orange Small offers a better speed/quality trade-off; choose Base when you want a little more headroom than Small without the cost of the 2.1B teacher.
Usage
The easiest way to use this model is through the
pulpie
package:
pip install pulpie
from pulpie import Extractor
extractor = Extractor(model="orange-base")
result = extractor.extract(html)
print(result.markdown) # clean Markdownprint(result.html) # clean HTMLprint(result.n_main, result.n_other) # blocks kept vs dropped
Extractor
auto-detects CUDA, Apple MPS, then CPU. See the
GitHub README
for batch and multi-GPU usage.
How it works
Pulpie runs a four-stage pipeline:
Simplify
— remove scripts, styles, and formatting noise; tag each block with a unique ID.
Chunk
— pack blocks into sequences of up to 8,192 tokens separated by
<|sep|>
markers (~80% of pages fit in one chunk).
Classify
— a single encoder forward pass labels every block (at its
<|sep|>
position) as content or boilerplate.
Reconstruct
— return the kept blocks as HTML, or convert to Markdown.
This model is a token-classification head over
EuroBERT-610m
, distilled from the 2.1B
Pulpie Orange Large
teacher (KL-divergence 0.7 + hard-label cross-entropy 0.3, temperature 2.0).
Benchmarks
WebMainBench, English subset (6,647 pages), ROUGE-5 F1:
Pulpie builds directly on the work of the MinerU-HTML and Dripper team (Ma et al., 2025). Their
simplify_html
preprocessing, block-level annotation scheme, and the WebMainBench benchmark are foundational to this work. Built on
EuroBERT
(Boizard et al., 2025).
Citation
@note{pulpie2026,
title = {Pulpie: Pareto-Optimal Models for Cleaning the Web},
author = {Minhas, Bhavnick and Nigam, Shreyash and Feyn Research},
year = {2026},
venue = {Feyn Field Notes}
}
Built by
Feyn
. Model weights and the
pulpie
library are licensed under Apache 2.0.
Runs of feyninc pulpie-orange-base on huggingface.co
50
Total runs
-18
24-hour runs
-29
3-day runs
-32
7-day runs
-20
30-day runs
More Information About pulpie-orange-base huggingface.co Model
pulpie-orange-base huggingface.co is an AI model on huggingface.co that provides pulpie-orange-base's model effect (), which can be used instantly with this feyninc pulpie-orange-base model. huggingface.co supports a free trial of the pulpie-orange-base model, and also provides paid use of the pulpie-orange-base. Support call pulpie-orange-base model through api, including Node.js, Python, http.
pulpie-orange-base huggingface.co is an online trial and call api platform, which integrates pulpie-orange-base's modeling effects, including api services, and provides a free online trial of pulpie-orange-base, you can try pulpie-orange-base online for free by clicking the link below.
feyninc pulpie-orange-base online free url in huggingface.co:
pulpie-orange-base is an open source model from GitHub that offers a free installation service, and any user can find pulpie-orange-base on GitHub to install. At the same time, huggingface.co provides the effect of pulpie-orange-base install, users can directly use pulpie-orange-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.