casehold / legalbert

huggingface.co
Total runs: 931
24-hour runs: 37
7-day runs: 223
30-day runs: 529
Model's Last Updated: July 02 2021
fill-mask

Introduction of legalbert

Model Details of legalbert

Legal-BERT

Model and tokenizer files for Legal-BERT model from When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset of 53,000+ Legal Holdings .

Training Data

The pretraining corpus was constructed by ingesting the entire Harvard Law case corpus from 1965 to the present ( https://case.law/ ). The size of this corpus (37GB) is substantial, representing 3,446,187 legal decisions across all federal and state courts, and is larger than the size of the BookCorpus/Wikipedia corpus originally used to train BERT (15GB).

Training Objective

This model is initialized with the base BERT model (uncased, 110M parameters), bert-base-uncased , and trained for an additional 1M steps on the MLM and NSP objective, with tokenization and sentence segmentation adapted for legal text (cf. the paper).

Usage

Please see the casehold repository for scripts that support computing pretrain loss and finetuning on Legal-BERT for classification and multiple choice tasks described in the paper: Overruling, Terms of Service, CaseHOLD.

Citation
@inproceedings{zhengguha2021,
        title={When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset},
        author={Lucia Zheng and Neel Guha and Brandon R. Anderson and Peter Henderson and Daniel E. Ho},
        year={2021},
        eprint={2104.08671},
        archivePrefix={arXiv},
        primaryClass={cs.CL},
        booktitle={Proceedings of the 18th International Conference on Artificial Intelligence and Law},
        publisher={Association for Computing Machinery}
}

Lucia Zheng, Neel Guha, Brandon R. Anderson, Peter Henderson, and Daniel E. Ho. 2021. When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset. In Proceedings of the 18th International Conference on Artificial Intelligence and Law (ICAIL '21) , June 21-25, 2021, São Paulo, Brazil. ACM Inc., New York, NY, (in press). arXiv: 2104.08671 [cs.CL] .

Runs of casehold legalbert on huggingface.co

931
Total runs
37
24-hour runs
145
3-day runs
223
7-day runs
529
30-day runs

More Information About legalbert huggingface.co Model

legalbert huggingface.co

legalbert huggingface.co is an AI model on huggingface.co that provides legalbert's model effect (), which can be used instantly with this casehold legalbert model. huggingface.co supports a free trial of the legalbert model, and also provides paid use of the legalbert. Support call legalbert model through api, including Node.js, Python, http.

casehold legalbert online free

legalbert huggingface.co is an online trial and call api platform, which integrates legalbert's modeling effects, including api services, and provides a free online trial of legalbert, you can try legalbert online for free by clicking the link below.

casehold legalbert online free url in huggingface.co:

https://huggingface.co/casehold/legalbert

legalbert install

legalbert is an open source model from GitHub that offers a free installation service, and any user can find legalbert on GitHub to install. At the same time, huggingface.co provides the effect of legalbert install, users can directly use legalbert installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

legalbert install url in huggingface.co:

https://huggingface.co/casehold/legalbert

Url of legalbert

legalbert huggingface.co Url

Provider of legalbert huggingface.co

casehold
ORGANIZATIONS

Other API from casehold