davanstrien / finepdfs-edu-purpose-classifier

huggingface.co
Total runs: 45
24-hour runs: 0
7-day runs: 22
30-day runs: 45
Model's Last Updated: September 10 2026
text-classification

Introduction of finepdfs-edu-purpose-classifier

Model Details of finepdfs-edu-purpose-classifier

FinePDFs-Edu document-purpose classifier

A proof of concept for using an agent to build a small classifier for data curation , trained and run with Hugging Face Jobs. This is the exact SetFit checkpoint used to classify 191,724 eligible documents from a 1% English FinePDFs-Edu sample.

Blog post · Labelled dataset · SetFit Jobs training recipe

Use on a text file
hf download davanstrien/finepdfs-edu-purpose-classifier predict.py --local-dir .
uv run --python 3.12 predict.py document.txt

Add --device cuda to use a CUDA GPU. The helper applies the same text preparation as the reported run. It prints the predicted purpose and scores in label order. Scores have not been calibrated as probabilities of correctness.

The model body is the 149M-parameter nomic-ai/modernbert-embed-base , fine-tuned with SetFit for 180 steps on 200 agent-labelled examples under a human-reviewed purpose guide, with a balanced logistic-regression head ( C=0.01 ). Seed 42 was the predefined model for reuse. Its six labels are:

  • administration_policy
  • exercise_assessment
  • instruction_reference
  • news_promotion
  • other
  • research_analysis
Input preparation matters

Use the pinned Granite tokenizer in predict.py to select excerpts: retain full text up to 990 preparation tokens, otherwise take 330 tokens each from the beginning, middle and end, separated by \n[...]\n . Prefix the result with classification: . The ModernBERT encoder then uses its own saved tokenizer with a maximum sequence length of 1,024. Its saved normalization module is retained; no additional encode-time normalization or automatic prompt is applied.

The helper accepts document text. It does not apply the dataset run's language and minimum-length eligibility checks. The reported dataset run checked document/page languages and excluded short or invalid records before inference.

Evaluation and limitations

On 120 source-separated, held-out agent-labelled documents: 65.8% accuracy; 0.556 macro-F1 . These are preliminary results against agent references, not independent human ground truth or an estimate for the entire corpus. There were no other predictions in that evaluation. The label guide has ambiguous boundaries, and excerpting can miss a document's dominant purpose.

These labels describe document function, not quality, factual reliability or retrieval relevance. English only was evaluated. Downstream retrieval/pretraining gains have not been tested. Review predictions and retain random samples for coverage before relying on a category.

Run record

Original inference script · Runtime configuration · Sampling manifest · Checkpoint provenance

The reproduction files are an unchanged record of the measured run, including the original private experiment-repository paths and pinned checksums. They are source material for inspection/adaptation, not a drop-in command targeting a reader's account. Use predict.py above to apply this public checkpoint to your own text; use the Jobs recipe to train a classifier for your own labels.

The 1% run took about 42 minutes on A10G-small and cost about $0.70 in running compute. Training and model-selection experiments cost approximately $2.90 separately. Agent and storage costs are excluded.

Runs of davanstrien finepdfs-edu-purpose-classifier on huggingface.co

45
Total runs
0
24-hour runs
0
3-day runs
22
7-day runs
45
30-day runs

More Information About finepdfs-edu-purpose-classifier huggingface.co Model

More finepdfs-edu-purpose-classifier license Visit here:

https://choosealicense.com/licenses/apache-2.0

finepdfs-edu-purpose-classifier huggingface.co

finepdfs-edu-purpose-classifier huggingface.co is an AI model on huggingface.co that provides finepdfs-edu-purpose-classifier's model effect (), which can be used instantly with this davanstrien finepdfs-edu-purpose-classifier model. huggingface.co supports a free trial of the finepdfs-edu-purpose-classifier model, and also provides paid use of the finepdfs-edu-purpose-classifier. Support call finepdfs-edu-purpose-classifier model through api, including Node.js, Python, http.

finepdfs-edu-purpose-classifier huggingface.co Url

https://huggingface.co/davanstrien/finepdfs-edu-purpose-classifier

davanstrien finepdfs-edu-purpose-classifier online free

finepdfs-edu-purpose-classifier huggingface.co is an online trial and call api platform, which integrates finepdfs-edu-purpose-classifier's modeling effects, including api services, and provides a free online trial of finepdfs-edu-purpose-classifier, you can try finepdfs-edu-purpose-classifier online for free by clicking the link below.

davanstrien finepdfs-edu-purpose-classifier online free url in huggingface.co:

https://huggingface.co/davanstrien/finepdfs-edu-purpose-classifier

finepdfs-edu-purpose-classifier install

finepdfs-edu-purpose-classifier is an open source model from GitHub that offers a free installation service, and any user can find finepdfs-edu-purpose-classifier on GitHub to install. At the same time, huggingface.co provides the effect of finepdfs-edu-purpose-classifier install, users can directly use finepdfs-edu-purpose-classifier installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

finepdfs-edu-purpose-classifier install url in huggingface.co:

https://huggingface.co/davanstrien/finepdfs-edu-purpose-classifier

Url of finepdfs-edu-purpose-classifier

finepdfs-edu-purpose-classifier huggingface.co Url

Provider of finepdfs-edu-purpose-classifier huggingface.co

davanstrien
ORGANIZATIONS

Other API from davanstrien

huggingface.co

Total runs: 13
Run Growth: 7
Growth Rate: 53.85%
Updated:September 24 2024
huggingface.co

Total runs: 12
Run Growth: 2
Growth Rate: 16.67%
Updated:April 11 2024