To do this, we are discovering and assembling the knowledge required to build artificial intelligence and returning it to the world.
Right now, we're training a 535 billion parameter mixture-of-experts model with 23 billion active parameters, on 18 trillion tokens of data. You can follow along here.
Open means that we share everything:
Selected writing: 8B dense retro 32B dense retro Delphi scaling suite Cluster scheduling with Iris LLM pretraining efficiency MoE quantile balancing
◇