GPT-NeoX

This repository records EleutherAI's library for training large-scale language models on GPUs. Our current framework is based on NVIDIA's Megatron Language Model and has been augmented with techniques from DeepSpeed as well as some novel optimizations. We aim to make this repo a centralized and accessible place to gather techniques for training large-scale autoregressive language models, and accelerate research into large-scale training.

For those looking for a TPU-centric codebase, we recommend Mesh Transformer JAX.

If you are not looking to train models with billions of parameters from scratch, this is likely the wrong library to use. For generic inference needs, we recommend you use the Hugging Face transformers library instead which supports GPT-NeoX models.

Project Samples

Project Activity

See All Activity >

License

Apache License V2.0

Follow GPT-NeoX

GPT-NeoX Web Site

Other Useful Business Software

Our Free Plans just got better! | Auth0

With up to 25k MAUs and unlimited Okta connections, our Free Plan lets you focus on what you do best—building great apps.

You asked, we delivered! Auth0 is excited to expand our Free and Paid plans to include more options so you can focus on building, deploying, and scaling applications without having to worry about your security. Auth0 now, thank yourself later.

Try free now

Rate This Project

User Reviews

Be the first to post a review of GPT-NeoX!

Additional Project Details

Programming Language

Python

Related Categories

Python Artificial Intelligence Software, Python Large Language Models (LLM), Python Deep Learning Frameworks, Python ChatGPT Apps, Python Generative AI

Registered

2023-03-23

Similar Business Software

GPT-NeoX

An implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library. This repository records EleutherAI's library for training large-scale language models on GPUs. Our current framework is based on NVIDIA's Megatron Language Model and has been augmented...

See Software
Alpa

Alpa aims to automate large-scale distributed training and serving with just a few lines of code. Alpa was initially developed by folks in the Sky Lab, UC Berkeley. Some advanced techniques used in Alpa have been written in a paper published in OSDI'2022. Alpa community is growing with new...

See Software
DeepSpeed

DeepSpeed is an open source deep learning optimization library for PyTorch. It's designed to reduce computing power and memory use, and to train large distributed models with better parallelism on existing computer hardware. DeepSpeed is optimized for low latency, high throughput...

See Software

Report inappropriate content

GPT-NeoX

Implementation of model parallel autoregressive transformers on GPUs

Get an email when there's a new version of GPT-NeoX

Project Samples

Project Activity

Categories

License

Follow GPT-NeoX

User Reviews

Additional Project Details

Programming Language

Related Categories

Registered