The pretraining corpus was constructed by ingesting the entire Harvard Law case corpus from 1965 to the present (
https://case.law/
). The size of this corpus (37GB) is substantial, representing 3,446,187 legal decisions across all federal and state courts, and is larger than the size of the BookCorpus/Wikipedia corpus originally used to train BERT (15GB).
Training Objective
This model is pretrained from scratch for 2M steps on the MLM and NSP objective, with tokenization and sentence segmentation adapted for legal text (cf. the paper).
The model also uses a custom domain-specific legal vocabulary. The vocabulary set is constructed using
SentencePiece
on a subsample (approx. 13M) of sentences from our pretraining corpus, with the number of tokens fixed to 32,000.
Usage
Please see the
casehold repository
for scripts that support computing pretrain loss and finetuning on Custom Legal-BERT for classification and multiple choice tasks described in the paper: Overruling, Terms of Service, CaseHOLD.
Citation
@inproceedings{zhengguha2021,
title={When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset},
author={Lucia Zheng and Neel Guha and Brandon R. Anderson and Peter Henderson and Daniel E. Ho},
year={2021},
eprint={2104.08671},
archivePrefix={arXiv},
primaryClass={cs.CL},
booktitle={Proceedings of the 18th International Conference on Artificial Intelligence and Law},
publisher={Association for Computing Machinery}
}
Lucia Zheng, Neel Guha, Brandon R. Anderson, Peter Henderson, and Daniel E. Ho. 2021. When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset. In
Proceedings of the 18th International Conference on Artificial Intelligence and Law (ICAIL '21)
, June 21-25, 2021, São Paulo, Brazil. ACM Inc., New York, NY, (in press). arXiv:
2104.08671 [cs.CL]
.
Runs of casehold custom-legalbert on huggingface.co
3.0K
Total runs
-49
24-hour runs
-385
3-day runs
-788
7-day runs
-2.4K
30-day runs
More Information About custom-legalbert huggingface.co Model
custom-legalbert huggingface.co
custom-legalbert huggingface.co is an AI model on huggingface.co that provides custom-legalbert's model effect (), which can be used instantly with this casehold custom-legalbert model. huggingface.co supports a free trial of the custom-legalbert model, and also provides paid use of the custom-legalbert. Support call custom-legalbert model through api, including Node.js, Python, http.
custom-legalbert huggingface.co is an online trial and call api platform, which integrates custom-legalbert's modeling effects, including api services, and provides a free online trial of custom-legalbert, you can try custom-legalbert online for free by clicking the link below.
casehold custom-legalbert online free url in huggingface.co:
custom-legalbert is an open source model from GitHub that offers a free installation service, and any user can find custom-legalbert on GitHub to install. At the same time, huggingface.co provides the effect of custom-legalbert install, users can directly use custom-legalbert installed effect in huggingface.co for debugging and trial. It also supports api for free installation.