flax-community / gpt-neo-125M-code-clippy-dedup

huggingface.co
Total runs: 22
24-hour runs: 0
7-day runs: 4
30-day runs: 14
Model's Last Updated: July 26 2021
text-generation

Introduction of gpt-neo-125M-code-clippy-dedup

Model Details of gpt-neo-125M-code-clippy-dedup

GPT-Neo-125M-Code-Clippy-Dedup

Please refer to our new GitHub Wiki which documents our efforts in detail in creating the open source version of GitHub Copilot

Model Description

PT-Neo-125M-Code-Clippy-Dedup is a GPT-Neo-125M model finetuned using causal language modeling on our deduplicated version of the Code Clippy Data dataset, which was scraped from public Github repositories (more information in the provided link). This model is specialized to autocomplete methods in multiple programming languages.

Training data

Code Clippy Data dataset .

Training procedure

In this model's training we tried to stabilize the training by limiting the types of files we were using to train to only those that contained file extensions for popular programming languages as our dataset contains other types of files as well such as .txt or project configuration files. We used the following extensions to filter by:

The training script used to train this model can be found here .

./run_clm_streaming_filter_flax.py \
    --output_dir $HOME/gpt-neo-125M-code-clippy-dedup \
    --model_name_or_path="EleutherAI/gpt-neo-125M" \
    --dataset_name $HOME/gpt-code-clippy/data_processing/code_clippy_filter.py \
    --data_dir $HOME/code_clippy_data/code_clippy_dedup_data \
    --text_column_name="text" \
    --do_train --do_eval \
    --block_size="2048" \
    --per_device_train_batch_size="8" \
    --per_device_eval_batch_size="16" \
    --preprocessing_num_workers="8" \
    --learning_rate="1e-4" \
    --max_steps 100000 \
    --warmup_steps 2000 \
    --decay_steps 30000 \
    --adam_beta1="0.9" \
    --adam_beta2="0.95" \
    --weight_decay="0.1" \
    --overwrite_output_dir \
    --logging_steps="25" \
    --eval_steps="500" \
    --push_to_hub="False" \
    --report_to="all" \
    --dtype="bfloat16" \
    --skip_memory_metrics="True" \
    --save_steps="500" \
    --save_total_limit 10 \
    --gradient_accumulation_steps 16 \
    --report_to="wandb" \
    --run_name="gpt-neo-125M-code-clippy-dedup-filtered-no-resize-2048bs" \
    --max_eval_samples 2000 \
    --save_optimizer true
Intended Use and Limitations

The model is finetuned text file from github repositories (mostly programming languages but also markdown and other project related files).

How to use

You can use this model directly with a pipeline for text generation. This example generates a different sequence each time it's run:


from transformers import AutoModelForCausalLM, AutoTokenizer, FlaxAutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained("flax-community/gpt-neo-125M-code-clippy-dedup")

tokenizer = AutoTokenizer.from_pretrained("flax-community/gpt-neo-125M-code-clippy-dedup")

prompt = """def greet(name):
  '''A function to greet user. Given a user name it should say hello'''
""" 

input_ids = tokenizer(prompt, return_tensors='pt').input_ids.to(device)

start = input_ids.size(1)

out = model.generate(input_ids, do_sample=True, max_length=50, num_beams=2, 

                     early_stopping=True, eos_token_id=tokenizer.eos_token_id, )

print(tokenizer.decode(out[0][start:]))
Limitations and Biases

The model is intended to be used for research purposes and comes with no guarantees of quality of generated code.

The paper "Evaluating Large Language Models Trained on Code" from OpenAI has a good discussion on what the impact of a large language model trained on code could be. Therefore, some parts of their discuss are highlighted here as it pertains to this dataset and models that may be trained from it. As well as some differences in views from the paper, particularly around legal implications .

  1. Over-reliance: This model may generate plausible solutions that may appear correct, but are not necessarily the correct solution. Not properly evaluating the generated code may cause have negative consequences such as the introduction of bugs, or the introduction of security vulnerabilities. Therefore, it is important that users are aware of the limitations and potential negative consequences of using this language model.
  2. Economic and labor market impacts: Large language models trained on large code datasets such as this one that are capable of generating high-quality code have the potential to automate part of the software development process. This may negatively impact software developers. However, as discussed in the paper, as shown in the Summary Report of software developers from O*NET OnLine , developers don't just write software.
  3. Security implications: No filtering or checking of vulnerabilities or buggy code was performed on the datase this model is trained on. This means that the dataset may contain code that may be malicious or contain vulnerabilities. Therefore, this model may generate vulnerable, buggy, or malicious code. In safety critical software, this could lead to software that may work improperly and could result in serious consequences depending on the software. Additionally, this model may be able to be used to generate malicious code on purpose in order to perform ransomware or other such attacks.
  4. Legal implications: No filtering was performed on licensed code. This means that the dataset may contain restrictive licensed code. As discussed in the paper, public Github repositories may fall under "fair use." However, there has been little to no previous cases of such usages of licensed publicly available code. Therefore, any code generated with this model may be required to obey license terms that align with the software it was trained on such as GPL-3.0. It is unclear the legal ramifications of using a language model trained on this dataset.
  5. Biases: The programming languages most represented in the dataset this model was trained on are Javascript and Python. Therefore, other, still popular languages such as C and C++, are less represented and therefore the models performance for these languages will be less comparatively. Additionally, this dataset only contains public repositories and so the model may not generate code that is representative of code written by private developers. No filtering was performed for potential racist, offensive, or otherwise inappropriate content. Therefore, this model may reflect such biases in its generation.

GPT-Neo-125M-Code-Clippy-Dedup is finetuned from GPT-Neo and might have inherited biases and limitations from it. See GPT-Neo model card for details.

Eval results

Coming soon...

Runs of flax-community gpt-neo-125M-code-clippy-dedup on huggingface.co

22
Total runs
0
24-hour runs
1
3-day runs
4
7-day runs
14
30-day runs

More Information About gpt-neo-125M-code-clippy-dedup huggingface.co Model

gpt-neo-125M-code-clippy-dedup huggingface.co

gpt-neo-125M-code-clippy-dedup huggingface.co is an AI model on huggingface.co that provides gpt-neo-125M-code-clippy-dedup's model effect (), which can be used instantly with this flax-community gpt-neo-125M-code-clippy-dedup model. huggingface.co supports a free trial of the gpt-neo-125M-code-clippy-dedup model, and also provides paid use of the gpt-neo-125M-code-clippy-dedup. Support call gpt-neo-125M-code-clippy-dedup model through api, including Node.js, Python, http.

gpt-neo-125M-code-clippy-dedup huggingface.co Url

https://huggingface.co/flax-community/gpt-neo-125M-code-clippy-dedup

flax-community gpt-neo-125M-code-clippy-dedup online free

gpt-neo-125M-code-clippy-dedup huggingface.co is an online trial and call api platform, which integrates gpt-neo-125M-code-clippy-dedup's modeling effects, including api services, and provides a free online trial of gpt-neo-125M-code-clippy-dedup, you can try gpt-neo-125M-code-clippy-dedup online for free by clicking the link below.

flax-community gpt-neo-125M-code-clippy-dedup online free url in huggingface.co:

https://huggingface.co/flax-community/gpt-neo-125M-code-clippy-dedup

gpt-neo-125M-code-clippy-dedup install

gpt-neo-125M-code-clippy-dedup is an open source model from GitHub that offers a free installation service, and any user can find gpt-neo-125M-code-clippy-dedup on GitHub to install. At the same time, huggingface.co provides the effect of gpt-neo-125M-code-clippy-dedup install, users can directly use gpt-neo-125M-code-clippy-dedup installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

gpt-neo-125M-code-clippy-dedup install url in huggingface.co:

https://huggingface.co/flax-community/gpt-neo-125M-code-clippy-dedup

Url of gpt-neo-125M-code-clippy-dedup

gpt-neo-125M-code-clippy-dedup huggingface.co Url

Provider of gpt-neo-125M-code-clippy-dedup huggingface.co

flax-community
ORGANIZATIONS

Other API from flax-community