Points represent fictional recruitment difficulty and opportunity cost. Each card uses 1 to 5 of your 20 game points.
A different pick.
A little more room.
Spend 5 points on one seat and you have 15 for the other six. Choose a 2-point card instead and you have three points more to work with.
Use those points elsewhere, or finish with them unspent. The choice is yours.
During the draft, a pick is unavailable if it would leave too few points to fill your remaining seats. You can still see that card and why it won’t fit.
All seven seats are filled. Spending all 20 points is a valid finish too.
5 points does not mean a better scientist.
A 5-point researcher is not “better” than a 1-point researcher. They are simply harder or more expensive to recruit in the game’s fictional frontier-AI talent market. Points do not measure scientific ability, quality, prestige or predicted future impact.
Building tools: Created, co-created, maintained or led delivery of reusable software, research tools, datasets, benchmarks or developer-facing products.
Open work: Publicly shared research, code, tools, data or technical writing; not necessarily open model weights.
The creator sets points by hand for a fictional frontier-AI talent market. They represent recruitment difficulty and opportunity cost, imagining factors such as current compensation and equity, institutional position, competing demand, scarcity and switching friction. Research interests, mission, autonomy and willingness to move matter too. These are provisional game settings, not measured recruitment offers or literal salary estimates.
Someone’s research interests or desire for autonomy might matter more than money. Their points are not a claim about their actual pay, availability or willingness to join your lab. These values are set by hand, not calculated from current compensation or recruitment data.
The current role pools do not all contain every point value. There are no measured percentiles behind these points. Roles are seats in this game, not claims about anyone's current job or availability.
The people are real. The lab is yours.
Choose seven people within 20 game points to make a lab card. Unspent points earn a Bootstrapped badge. Optional challenges check your team automatically. These rules do not predict what people would achieve together. Lab names and commentary are fictional team-level jokes.
Frontier Draft is an independent fan-made parody game. Appearing here does not imply participation, affiliation or endorsement. No one is assigned an individual personality score.
Something wrong, or prefer to be removed?
Contact @falseconjecture on X with the card name and the correction or removal request. For a factual correction, a public source helps. Feedback on the game rules is welcome too.
Look up a card
108 peopleGame points, public work and primary sources for every person. These sources explain the work; they do not justify or calculate the fictional point values.
108 cards shown
- Research Lead
Aidan Gomez
3 ptsTransformer co-author, Cohere co-founder
Co-authored Attention Is All You Need as one of the eight researchers who introduced the Transformer architecture for sequence tasks. Co-founded Cohere in 2019, building a company around language models and tools for working with text, including generation and semantic search.
Attention Is All You NeedCohere: founders and company history - Research Lead
Albert Gu
3 ptsState space models and Mamba co-creator
With Tri Dao, wrote Mamba, introducing selective state-space models that decide which information to retain as they process a sequence. Co-created the accompanying Mamba library, releasing code for the model and its selective state-space operations so others could run and study the approach.
Mamba: Linear-Time Sequence Modeling with Selective State SpacesMamba — reusable state-space library - Research Lead
Alec Radford
5 ptsFirst author of GPT, GPT-2, and CLIP
Was first author of CLIP, training image and text representations together so a model could recognize categories described in words. Co-authored Proximal Policy Optimization, a reinforcement-learning method that reuses collected experience for several training updates and was tested on games and simulated movement.
Proximal Policy Optimization AlgorithmsCLIP: Learning Transferable Visual Models From Natural Language Supervision - Wildcard only
Alexandr Wang
3 ptsFounded Scale AI, superintelligence-lab leadership
Co-founded Scale to provide training data for machine-learning teams. Its early APIs handled image annotation, transcription and other data tasks. The company served applications including self-driving vehicles, intelligent assistants and voice analytics.
Scale — early training-data platform announcement - Safety Lead
Amanda Askell
2 ptsThe philosopher who shapes Claude's character
Co-authored A General Language Assistant as a Laboratory for Alignment, using an assistant model to study how training changes its responses. The experiments compared prompting with learning from preferences, examining how each approach affected helpfulness, honesty and harmlessness in the assistant's behavior.
A General Language Assistant as a Laboratory for Alignment - Research Lead
Andrej Karpathy
5 ptsOpenAI founding member, ex-Tesla Autopilot lead, the internet's AI teacher
Co-authored the ImageNet challenge report, documenting how image recognition systems were evaluated on classification and object detection. Created Neural Networks: Zero to Hero, a course with videos and notebooks that build neural networks in code, from a small automatic-differentiation engine to language models.
ImageNet Large Scale Visual Recognition ChallengeNeural Networks: Zero to Hero - Chief Scientist
Andrew Chi-Chih Yao
4 ptsTuring Award, founded Tsinghua's Yao Class, trained a generation
Developed research in computational complexity and cryptography, including work on communication complexity and secure computation: how parties can calculate together without revealing all their inputs. Founded Tsinghua's Yao Class in 2005, creating an undergraduate program focused on computer science.
Andrew Chi-Chih Yao — TsinghuaAndrew Yao: research and the Yao Class - Wildcard only
Andrew Ng
3 ptsTaught everyone's first ML course, serial founder
Worked on Stanford's autonomous helicopter project, applying reinforcement learning to flight control. The project studied how learning algorithms could produce control policies for demanding maneuvers on a real helicopter, taking questions about an agent's behavior beyond simulations and into physical flight.
Stanford autonomous helicopter project - Infra Lead
Andrew Tulloch
2 ptsML systems veteran across frontier labs
Worked on machine-learning systems at Meta and contributed to training GPT-4o, GPT-4.5 and o3 at OpenAI, as recorded in his own biography. Studied mathematics at Sydney and Cambridge before his work on the systems used to train large models.
Andrew Tulloch — personal biography - Wildcard only
Arthur Mensch
3 ptsMistral co-founder, Europe's frontier bet
Co-authored the Mistral 7B report, describing a language model built with grouped-query and sliding-window attention. The team released a base model and an instruction-following version. The report explains the architecture choices behind a model intended to use compute and memory efficiently.
Mistral 7B - Evals Lead
Arvind Narayanan
3 ptsCo-wrote AI Snake Oil and a Bitcoin/cryptocurrency textbook
Worked on Princeton's Web Transparency and Accountability Project, studying how companies collect and use people's information. His research has also examined cultural stereotypes in machine learning. Co-wrote AI Snake Oil to help readers assess what different kinds of AI can actually do.
AI Snake Oil — author biographies - Research Lead
Ashish Vaswani
3 ptsFirst author of Attention Is All You Need
Co-authored the 2017 Transformer paper with seven collaborators, introducing a sequence model driven by attention. The team built an encoder and decoder around attention rather than recurrence, then tested the architecture on machine translation and reported how it changed training efficiency.
Attention Is All You Need - Post-training Lead
Barret Zoph
3 ptsNeural architecture search, led post-training at OpenAI
With Quoc Le, developed neural architecture search, training a controller to propose network designs and improve them using measured validation results. Co-authored Switch Transformers with William Fedus and Noam Shazeer, routing each token to a selected expert network to limit the computation used.
Neural Architecture Search with Reinforcement LearningSwitch Transformers - Evals Lead
Beth Barnes
3 ptsFounded METR, tests whether models can scheme
Co-authored METR's randomized study of experienced open-source developers using early-2025 AI tools. The study compared task completion times with and without those tools in repositories the developers already knew, testing productivity claims against work in familiar, existing codebases.
AI and experienced developer productivity — METR - Safety Lead
Buck Shlegeris
2 ptsCo-founded Redwood Research, works on the AI control agenda
Co-authored AI Control, a study that tests safeguards while assuming a model may deliberately try to get around them. The team evaluated protocols for getting useful work from an untrusted model, measuring how well the safeguards held up against attempts to undermine them.
AI Control: Improving Safety Despite Intentional Subversion - Infra Lead
Chris Lattner
3 ptsCreated LLVM and Swift, building Mojo
Co-created LLVM and co-wrote its compiler-framework paper with Vikram Adve, using a shared intermediate representation to analyze and transform programs. Started and led Swift's language design and implementation, also helping design its compiler architecture and interactive programming tools.
LLVM compilation framework paperChris Lattner: LLVM and Swift contributions - Chief Scientist
Chris Manning
2 ptsCo-created GloVe, taught Stanford's CS224N, ACL president 2015
Developed GloVe with Jeffrey Pennington and Richard Socher. The project released code and pretrained word vectors learned from word co-occurrence counts. Taught Stanford's CS224N course on natural language processing with deep learning, with public lectures and assignments covering how neural models work with language.
GloVe — Stanford NLPStanford CS224N: 2019 course archive - Safety Lead
Chris Olah
4 ptsFoundational mechanistic interpretability work; Anthropic co-founder
Co-authored Distill's Feature Visualization work, which generates images to investigate the patterns that activate parts of a neural network. Co-authored the Circuits research series, tracing connections between neurons to study how simpler visual features combine into more complex behavior.
Feature Visualization — DistillAn Introduction to Circuits — Distill - Wildcard only
Clem Delangue
3 ptsCo-founded Hugging Face, testified to the US Congress on open AI
Gave congressional testimony in 2023 about Hugging Face's platform for sharing models and datasets. He described its approach to open science and argued for lowering barriers to participation in AI development, explaining why researchers and developers outside a few large organizations should be able to contribute.
Clement Delangue — congressional testimony - Evals Lead
Clémentine Fourrier
2 ptsCo-built the Open LLM Leaderboard, evaluation research
Helped maintain the Open LLM Leaderboard, a service for comparing models through a shared evaluation process. Co-authored Zephyr, which trained a smaller assistant using preferences supplied by a larger model. The team released its models and training recipe so others could study the approach.
Zephyr: Direct Distillation of LM AlignmentOpen LLM Leaderboard — Hugging Face - Evals Lead
Dan Hendrycks
4 ptsCreated MMLU, directs the Center for AI Safety
Co-authored MMLU, a benchmark that tests language models across 57 academic and professional subjects. The study examined differences in performance across subjects and whether models could recognize when they were wrong, going beyond whether a model answered individual questions correctly.
Measuring Massive Multitask Language Understanding - Wildcard only
Dario Amodei
4 ptsCo-authored GPT-3 and scaling laws, co-founded Anthropic
Co-authored Deep Reinforcement Learning from Human Preferences, studying how people could guide agents by comparing examples of their behavior. The team used those comparisons to train a reward model. This explored a way to specify goals that are hard to capture in a reward function written by hand.
Deep reinforcement learning from human preferences - Chief Scientist
David Silver
3 ptsLed AlphaGo and AlphaZero, 2019 ACM Prize in Computing
Led AlphaGo research and co-authored its original paper, combining neural networks with tree search to choose moves in Go. Co-authored AlphaZero, showing how one reinforcement-learning approach could learn chess, shogi and Go through self-play without training on expert games.
Mastering Chess and Shogi by Self-PlayAlphaGo: neural networks and tree search - Post-training Lead
Daya Guo
3 ptsFirst-named author on the DeepSeek-R1 reasoning paper
Was first-named author of the DeepSeek-R1 research, which investigated reinforcement learning for language-model reasoning. The team compared R1-Zero's reinforcement-learning-only approach with R1's pipeline using initial supervised examples, and published the training recipe and its reasoning evaluations.
DeepSeek-R1 - Chief Scientist
Demis Hassabis
5 ptsDeepMind co-founder, Nobel laureate for AlphaFold
Co-authored the original AlphaGo paper. The team combined neural networks, tree search and self-play to build a program that defeated professional Go player Fan Hui. Also co-authored the AlphaFold paper describing how the team predicted a protein's three-dimensional structure from its amino-acid sequence.
AlphaGo: neural networks and tree searchHighly accurate protein structure prediction with AlphaFold - Wildcard only
Dwarkesh Patel
2 ptsThe podcast where researchers confess their timelines
Publishes long interviews and written research questions about AI, alongside conversations about science and history. His Questions about the Future of AI essay draws out questions about changing capabilities and deployment, giving readers a way to follow the uncertainties behind the interview discussions.
Dwarkesh PodcastQuestions about the Future of AI - Infra Lead
Dylan Patel
2 ptsSemiAnalysis founder; semiconductor and AI infrastructure research
Founded SemiAnalysis, a research publication covering the chips and datacenters behind AI. Its reporting follows semiconductor supply chains and the costs of training and running models, connecting hardware decisions with the software workloads they need to support.
Dylan Patel — SemiAnalysisSemiAnalysis research coverage - Safety Lead
Eliezer Yudkowsky
3 ptsFounded MIRI, has been warning you since 2008
Wrote about logical uncertainty: how an agent should reason about a mathematical statement when it has not yet worked out the answer. His MIRI discussions connect that problem with artificial agents that must make decisions despite limits on the time and computation available for reasoning.
Yudkowsky on Logical Uncertainty — MIRI - Wildcard only
Elon Musk
5 ptsFounded xAI, co-founded OpenAI, left its board in 2018
Wrote Tesla's 2006 master plan, setting out a staged approach to developing electric cars. The plan began with an expensive sports car and described using that path to reach more affordable vehicles. It connected those product choices with a longer-term move toward sustainable energy.
The Secret Tesla Motors Master Plan - Safety Lead
Evan Hubinger
2 ptsAlignment stress-testing and deceptive alignment research
Co-authored Risks from Learned Optimization, examining whether training can produce a model that pursues an internal objective of its own. The paper distinguishes the objective used to train a model from the one its learned computations might pursue, and examines how those can come apart.
Risks from Learned Optimization in Advanced Machine Learning Systems - Chief Scientist
Fei-Fei Li
4 ptsCreated ImageNet, founded World Labs, spatial intelligence
Created ImageNet with collaborators, building an image dataset that researchers could use to train and compare visual recognition systems. Co-authored the ImageNet challenge report, documenting the benchmark's image classification and object detection tasks, evaluation methods and remaining difficulties.
ImageNet Large Scale Visual Recognition Challenge - Research Lead
François Chollet
4 ptsCreated Keras and the ARC-AGI benchmark, co-founded Ndea
Created Keras, a deep-learning API that lets developers assemble, train and reuse neural-network components. Designed the ARC benchmark and wrote On the Measure of Intelligence, using small visual puzzles to examine how a system learns unfamiliar tasks from a few examples.
On the Measure of IntelligenceKeras — project repository - Wildcard only
Gary Marcus
2 ptsFounded Geometric Intelligence, testified to the US Senate on AI
Wrote Deep Learning: A Critical Appraisal in 2018, setting out concerns about what deep learning methods could and could not do. The paper argues for combining those methods with other approaches. It separates success on particular benchmarks from the wider problem of building general intelligence.
Deep Learning: A Critical Appraisal - Chief Scientist
Geoffrey Hinton
5 ptsDeep learning pioneer, Turing Award, Nobel laureate
Co-authored AlexNet with Alex Krizhevsky and Ilya Sutskever, documenting how a deep convolutional network trained on GPUs could recognize objects in ImageNet. With Yann LeCun and Yoshua Bengio, wrote the 2015 Deep Learning review explaining learned representations, convolutional networks and recurrent networks.
Deep learning — NatureAlexNet: ImageNet Classification with Deep Convolutional Neural Networks - Infra Lead
Georgi Gerganov
3 ptsCreated llama.cpp, made local inference a movement
Created llama.cpp, a C and C++ project for running language models locally instead of relying on a hosted service. The project provides quantized model formats and support for different hardware, giving developers practical controls over memory use and where inference runs.
Georgi Gerganov — projectsllama.cpp — project repository - Infra Lead
Greg Brockman
4 ptsOpenAI co-founder and president, ships through the night
Co-founded OpenAI and co-wrote its 2015 introduction with Ilya Sutskever, setting out the new research organization's plans. Co-authored the OpenAI Five report, documenting the distributed training system and sustained self-play used to build a Dota 2 agent.
Dota 2 with Large Scale Deep Reinforcement LearningIntroducing OpenAI - Safety Lead
Helen Toner
3 ptsSat on OpenAI's nonprofit board through the 2023 leadership crisis
Co-authored CSET's Preparing for AI Agents report, examining systems that can take actions beyond a chat conversation. The report considers how those systems are developing and what risks they can amplify, then discusses places where technical safeguards or governance could intervene.
Preparing for AI Agents — CSET - Infra Lead
Horace He
1 ptPyTorch systems work; co-developed FlexAttention
Co-authored PyTorch's introduction to FlexAttention, which lets developers write their own attention patterns without starting from a custom GPU kernel. The interface compiles those patterns into efficient kernels. It gives model developers a way to experiment with attention while reusing the surrounding PyTorch tools.
FlexAttention — PyTorch - Post-training Lead
Hyung Won Chung
3 ptsInstruction tuning and reasoning research
Co-authored the Flan instruction-tuning study, testing how model size, task variety and examples containing reasoning steps affected the resulting models. The team released Flan-T5 checkpoints alongside the research, letting others run the instruction-tuned models and examine their behavior.
Scaling Instruction-Finetuned Language Models - Chief Scientist
Ilya Sutskever
5 ptsAlexNet co-author, OpenAI co-founder, founded Safe Superintelligence
Co-authored AlexNet with Alex Krizhevsky and Geoffrey Hinton. Their 2012 paper described a convolutional network trained on GPUs to recognize objects in ImageNet. With Oriol Vinyals and Quoc Le, co-wrote the 2014 sequence-to-sequence paper on using recurrent networks for translation.
Sequence to Sequence Learning with Neural NetworksAlexNet: ImageNet Classification with Deep Convolutional Neural Networks - Infra Lead
Ion Stoica
4 ptsCo-created Apache Spark and Ray, co-founded Databricks and Anyscale
Co-authored the Ray paper, describing a distributed framework for running parallel tasks and programs that keep their own state. The team tested that shared interface on reinforcement-learning workloads, where collecting experience, training models and running simulations place different demands on a cluster.
Ray: A Distributed Framework for Emerging AI Applications - Safety Lead
Jack Clark
3 ptsAnthropic co-founder, writes Import AI, co-chaired the AI Index
Co-authored the GPT-3 report, which studied whether a language model could learn tasks from examples included in its prompt. The report also examined limitations and societal implications of large text-generation systems. His Import AI newsletter explains research and policy developments for a public audience.
Language Models are Few-Shot LearnersImport AI by Jack Clark - Evals Lead
Jacob Steinhardt
3 ptsForecasting and evaluation research, founded Transluce
Co-authored a mechanistic study of grokking in small Transformers trained on modular arithmetic. The team reconstructed the algorithm the models learned and used it to measure progress before strong test performance appeared, connecting an internal computation with the later change in generalization.
Progress measures for grokking via mechanistic interpretability - Evals Lead
Jaime Sevilla
3 ptsCo-founded Epoch AI, which publishes AI compute and trend data
Co-authored Compute Trends Across Three Eras of Machine Learning, collecting historical evidence about the resources used to train models. Co-founded Epoch AI, which began by organizing model and compute data. The work gives researchers a record for studying changes in training scale over time.
Compute Trends Across Three Eras of Machine LearningWhat is Epoch? — Epoch AI - Chief Scientist
Jakub Pachocki
3 ptsLed GPT-4 pretraining at OpenAI
Led the development of GPT-4 and OpenAI Five, according to OpenAI's account of his research contributions. Co-authored the OpenAI Five report, documenting how a distributed reinforcement-learning system learned Dota 2 through self-play despite hidden information, long games and many possible actions.
Dota 2 with Large Scale Deep Reinforcement LearningOpenAI on Jakub Pachocki's research contributions - Post-training Lead
Jan Leike
4 ptsRLHF pioneer, led superalignment research
Co-authored Deep Reinforcement Learning from Human Preferences, studying how people's comparisons of behavior clips could train a reward model. Later co-authored InstructGPT, applying human demonstrations and preferences over text responses to the problem of training language models to follow instructions.
Training language models to follow instructions with human feedbackDeep reinforcement learning from human preferences - Chief Scientist
Jared Kaplan
3 ptsScaling laws co-author, Anthropic co-founder
Co-authored Scaling Laws for Neural Language Models, measuring how model size, training data and compute relate to prediction loss and training-budget allocation. Also co-authored the GPT-3 paper, which tested whether examples in a prompt could specify tasks without updating the model's weights.
Scaling Laws for Neural Language ModelsLanguage Models are Few-Shot Learners - Research Lead
Jason Wei
3 ptsChain-of-thought and emergent abilities papers
Was first author of the chain-of-thought prompting study, testing examples that show intermediate reasoning steps on arithmetic, commonsense and symbolic tasks. Co-authored Emergent Abilities of Large Language Models, examining behaviors observed at larger model scales and whether smaller models' results predicted them.
Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsEmergent Abilities of Large Language Models - Chief Scientist
Jeff Dean
4 ptsCo-created MapReduce, Bigtable, TensorFlow, drove the TPU program
Co-created MapReduce with Sanjay Ghemawat. Their system let programmers process large datasets while its runtime handled how work was split, scheduled and recovered after machine failures. Co-authored the TensorFlow systems paper, describing software for running machine-learning computations across different devices and distributed machines.
MapReduce — Google ResearchTensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems - Wildcard only
Jensen Huang
5 ptsCo-founded NVIDIA in 1993, bet the company on CUDA
Co-founded NVIDIA in 1993, building a company around graphics computing. In public presentations on scientific computing, he describes GPU hardware and software libraries as parts of the same computing system, with uses that extend beyond training and running language models.
Jensen Huang — NVIDIA biographyScientific computing at SC24 — NVIDIA - Research Lead
Jie Tang
3 ptsLed development of Tsinghua's GLM and ChatGLM model family
Co-authored GLM, developing a language-model pretraining approach with collaborators that learns by generating missing spans of text. The team tested whether this shared training method could support both language understanding and text generation, instead of designing a separate pretraining setup for each task.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling - Research Lead
Jim Fan
2 ptsCreated Voyager, led NVIDIA's GR00T robot project
Co-authored Voyager under the name Linxi Fan, helping build a Minecraft agent that uses a language model to write executable skills. The team designed a growing skill library and feedback loop so the agent could revise its code and reuse what it had learned on later tasks.
Voyager: An Open-Ended Embodied Agent - Research Lead
Joelle Pineau
2 ptsRan Meta FAIR, open science champion
Co-authored the report on the NeurIPS 2019 reproducibility program, examining code submission, a community replication challenge and an experimental reporting checklist. The team documented what these efforts revealed and proposed ways for researchers and conferences to make machine-learning results easier to inspect and reproduce.
Improving Reproducibility in Machine Learning Research - Post-training Lead
John Schulman
5 ptsCo-developed PPO; InstructGPT co-author
Co-developed Proximal Policy Optimization, a reinforcement-learning method that allows several training updates from the same batch of experience. Co-authored InstructGPT, combining human-written examples with rankings of model responses to train language models to follow instructions more closely.
Proximal Policy Optimization AlgorithmsTraining language models to follow instructions with human feedback - Research Lead
Junyang Lin
3 ptsLed Alibaba's Qwen open-model team
Led work on the Qwen open-model family and co-authored its technical report, covering pretrained language models and versions trained to follow instructions. The team's release included chat and code-specialized models, with published evaluations on language, mathematics and programming tasks.
Qwen Technical Report - Wildcard only
Kai-Fu Lee
3 ptsWrote AI Superpowers, founded 01.AI
Wrote his Carnegie Mellon doctoral thesis on continuous speech recognition with a large vocabulary. The research addressed speaker independence: recognizing connected speech without building a separate system around one person's voice. It focused on continuous utterances rather than asking a speaker to pause between isolated words.
Kai-fu Lee — doctoral thesis, Carnegie Mellon - Chief Scientist
Kaiming He
3 ptsResNet lead author; residual learning for visual recognition
Was first author of ResNet, introducing residual connections with collaborators to make very deep neural networks easier to train for image recognition. Co-authored Mask R-CNN, extending object detection so the system could also mark the pixels belonging to each detected object.
Deep Residual Learning for Image RecognitionMask R-CNN - Post-training Lead
Lewis Tunstall
2 ptsHugging Face alignment work, co-wrote the NLP with Transformers book
Co-created the original Open LLM Leaderboard, giving people a shared place to compare openly available language models on published evaluations. Co-authored Zephyr, training chat models with AI-generated preference rankings and releasing the models, code and recipes for studying direct preference optimization.
Zephyr: Direct Distillation of LM AlignmentOpen LLM Leaderboard — Hugging Face - Post-training Lead
Liam Fedus
3 ptsSwitch Transformer co-author, co-founded Periodic Labs
Co-authored Switch Transformers with Barret Zoph and Noam Shazeer, using selected expert networks so a large model does not activate every parameter for every token. Publishing as William Fedus, also co-authored the study of emergent abilities in language models, examining behaviors observed as model scale increased.
Switch TransformersEmergent Abilities of Large Language Models - Wildcard only
Liang Wenfeng
4 ptsFounded DeepSeek, shocked the market on a budget
Co-authored the DeepSeek-V2 report, where he is listed as Wenfeng Liang. The model combines a mixture-of-experts design with attention that uses less memory. The report studies how those choices can reduce the resources needed to run a language model.
DeepSeek-V2 - Post-training Lead
Long Ouyang
3 ptsFirst author of InstructGPT, the paper that made models follow instructions
Was first author of InstructGPT, documenting a training pipeline that starts with human-written examples and then learns from people's rankings of model responses. The study compared instruction-tuned models with larger base models, measuring human preferences and reporting remaining problems with truthfulness and harmful outputs.
Training language models to follow instructions with human feedback - Research Lead
Lucas Beyer
3 ptsVision Transformer co-author, serial frontier-lab researcher
Co-authored the Vision Transformer paper, testing whether a Transformer could recognize images presented as sequences of small patches. With collaborators, developed SigLIP's training objective for matching images and text, and released models so others could use and evaluate the method.
An Image is Worth 16x16 WordsSigmoid Loss for Language Image Pre-Training - Post-training Lead
Luke Metz
2 ptsCo-founded Thinking Machines Lab, learned optimizers research
Co-authored research that trains neural networks to make optimizer updates, testing whether learning across many tasks could transfer to unseen problems. Was first author of VeLO, scaling this approach into an optimizer that reads gradients and produces parameter updates for another neural network.
Training more effective learned optimizersVeLO: Training Versatile Learned Optimizers by Scaling Up - Safety Lead
Max Tegmark
2 ptsFuture of Life Institute, the pause letter
Co-authored Towards Guaranteed Safe AI, a proposal for using formal verification to support safety claims about AI systems. The framework separates the model of the world from the safety requirements and the process used to check them, so each assumption can be examined.
Towards Guaranteed Safe AI - Evals Lead
Mike Knoop
3 ptsCo-founded Zapier, co-founded the ARC Prize Foundation
Co-founded Zapier, a platform for automating work between software applications. Co-founded the ARC Prize Foundation, which supports research through reasoning benchmarks and competitions. Its stated focus is how systems adapt to unfamiliar problems, rather than how well they reproduce solutions to familiar tasks.
ARC Prize Foundation — mission and teamMike Knoop — Zapier - Safety Lead
Miles Brundage
3 ptsLong-running work on AI policy and deployment readiness
Co-authored Toward Trustworthy AI Development, a report on making claims about AI systems easier to verify. The report proposes technical and institutional checks for developers' claims, asking what evidence would let other people assess safety rather than rely on the developer's assurances alone.
Toward Trustworthy AI Development - Wildcard only
Mira Murati
3 ptsEx-OpenAI CTO, founded Thinking Machines Lab
Co-founded Thinking Machines Lab, whose published mission includes making AI systems easier for people to customize. The company's founding description centers collaboration between people and machines, and explains its research and product work in terms of giving users more ways to guide a system.
Thinking Machines — founding team and missionThinking Machines — NVIDIA partnership announcement - Post-training Lead
Nathan Lambert
1 ptOpen model post-training; author of the RLHF Book
Wrote the openly available RLHF Book, explaining preference data, reward models and reinforcement learning for language-model post-training. Was first author of the Tulu 3 report, releasing post-training methods with collaborators along with models, training data and code for others to inspect and reuse.
RLHF Book — Nathan LambertTulu 3: Pushing Frontiers in Open Language Model Post-Training - Safety Lead
Neel Nanda
3 ptsMechanistic interpretability lead, prolific educator
Created TransformerLens, a library for inspecting the internal computations of Transformer language models. Co-authored a study that reconstructed how small Transformers learned modular arithmetic. The team used that mechanism to track learning before the models' test accuracy showed generalization.
Progress measures for grokking via mechanistic interpretabilityTransformerLens — interpretability library - Research Lead
Noam Brown
4 ptsSuperhuman poker and Diplomacy AI, reasoning model research
Created the poker systems Libratus and Pluribus with his Carnegie Mellon advisor, studying decisions in games where players cannot see each other's cards. Helped develop CICERO with teammates, combining strategic planning with natural-language negotiation to play the board game Diplomacy.
Noam Brown — Carnegie Mellon thesis proposalNoam Brown: poker and Diplomacy research - Research Lead
Noam Shazeer
4 ptsTransformer co-author, founded Character.AI
Co-authored the Transformer paper, replacing recurrent sequence processing with attention in a model demonstrated on machine translation. Was first author of the sparsely gated mixture-of-experts paper, developing a layer that selects a small number of expert networks for each input instead of using every expert.
Attention Is All You NeedOutrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer - Evals Lead
Ofir Press
2 ptsCo-created SWE-bench and SWE-agent for coding evaluation
Co-authored SWE-bench, a benchmark built around real issues in existing software repositories. The tasks ask a model to understand an issue and the code around it, then propose a change that is checked with tests. The benchmark uses actual repository maintenance problems.
SWE-bench: Can Language Models Resolve Real-World GitHub Issues? - Research Lead
Oriol Vinyals
4 ptsseq2seq and AlphaStar, Gemini co-lead
Co-wrote the sequence-to-sequence paper with Ilya Sutskever and Quoc Le, using recurrent networks to translate an input sentence into an output sentence. Also co-authored the AlphaFold paper, contributing to the team's published system for predicting protein structures from amino-acid sequences.
Sequence to Sequence Learning with Neural NetworksHighly accurate protein structure prediction with AlphaFold - Safety Lead
Paul Christiano
4 ptsCo-authored foundational work on learning from human preferences
Co-authored the 2017 paper Deep Reinforcement Learning from Human Preferences, an early study of training agents from people's feedback. The team asked people to compare short behavior clips. A learned reward model then used those comparisons to guide agents without a reward function written for each task.
Deep reinforcement learning from human preferences - Evals Lead
Percy Liang
4 ptsCo-created the SQuAD and HELM evaluation benchmarks
Co-authored HELM, an evaluation framework for comparing language models across different scenarios. The team measured accuracy alongside robustness, fairness and other properties. It made the evaluation methods explicit so readers could see which parts of model behavior a result did, and did not, measure.
Holistic Evaluation of Language Models - Research Lead
Quoc Le
3 ptsseq2seq and AutoML, longtime Google Brain
Co-wrote the sequence-to-sequence paper with Ilya Sutskever and Oriol Vinyals, using recurrent networks to learn translation from paired sentences. With Barret Zoph, developed neural architecture search, training a controller to propose network designs and learn from their measured performance.
Neural Architecture Search with Reinforcement LearningSequence to Sequence Learning with Neural Networks - Chief Scientist
Richard Sutton
4 ptsWrote the RL textbook and The Bitter Lesson, Turing Award
Co-wrote Reinforcement Learning: An Introduction with Andrew Barto, explaining how an agent can learn from rewards while interacting with an uncertain environment. The textbook develops the algorithms step by step, with examples connecting prediction and decision-making to games, psychology and neuroscience.
Reinforcement Learning — MIT Press - Evals Lead
Rishi Bommasani
2 ptsCo-authored the Foundation Model Transparency Index
Co-authored HELM, an evaluation project that publishes its methods and model outputs for other researchers to examine. The work compares models across multiple scenarios and measures, making it possible to inspect differences in behavior that a single headline benchmark result would leave out.
Holistic Evaluation of Language Models - Safety Lead
Rohin Shah
2 ptsCreated and wrote the Alignment Newsletter since 2018
Co-authored research on goal misgeneralization, where an agent learns a goal that works during training but causes problems in a new situation. The work examines how this can happen even when the training reward is correct, separating it from mistakes in writing the reward itself.
How undesired goals can arise with correct rewards — Google DeepMind - Post-training Lead
Ross Taylor
2 ptsLed Galactica, reasoning and post-training research
Led the Galactica project and co-authored its paper, building a language model trained on scientific and technical material. The team evaluated scientific question answering, mathematical writing and citation generation, publishing experiments that explored what a language model could and could not do with scientific text.
Galactica: A Large Language Model for Science - Wildcard only
Sam Altman
5 ptsCo-founded OpenAI, ran Y Combinator 2014-2019
Wrote the Startup Playbook, collecting advice for people building a company at an early stage. It covers choosing an idea and a team, building a product, and the work of execution. Hiring and fundraising are treated as practical parts of getting the company operating.
Startup Playbook — Sam Altman - Evals Lead
Sayash Kapoor
2 ptsCo-wrote AI Snake Oil, agent benchmark critique
Co-wrote AI Snake Oil, examining the different technologies grouped under the AI label and the claims made about them. The book and accompanying newsletter ask where evidence supports an application and where promotion outruns it, helping readers separate a useful demonstration from a claim that has not been tested.
AI Snake Oil — book announcement - Post-training Lead
Sebastian Raschka
2 ptsTeaches LLM training from scratch to a very large audience
Wrote Build a Large Language Model (From Scratch), teaching readers to implement, pretrain and fine-tune a small GPT-style model in PyTorch. Published the book's code and notebooks in LLMs-from-scratch, letting readers run each stage and inspect the training steps themselves.
LLMs from scratch — official repositoryBuild a Large Language Model (From Scratch): publisher's page - Evals Lead
Sébastien Bubeck
3 ptsSparks of AGI paper, small-model research
Co-authored Sparks of Artificial General Intelligence, an exploratory study of an early GPT-4 model on tasks including mathematics and coding. The paper records capabilities and limitations, then argues for an interpretation of those results. Its claims are the authors' assessment, not a settled definition of intelligence.
Sparks of Artificial General Intelligence - Chief Scientist
Shane Legg
2 ptsDeepMind co-founder, helped popularize the term AGI
With Marcus Hutter, proposed a mathematical definition of machine intelligence based on how an agent performs across a broad range of environments. Co-authored Deep Reinforcement Learning from Human Preferences, studying how people's comparisons of short behavior clips could guide an agent's learning.
Universal Intelligence: A Definition of Machine IntelligenceDeep reinforcement learning from human preferences - Wildcard only
Simon Willison
1 ptDocuments the entire field in public, in real time
Created Datasette, a tool for exploring and publishing data, and co-created the Django web framework. Publishes programming experiments and explanations on his own site. The writing often accompanies reusable tools, giving readers both an account of the experiment and software they can try.
Simon Willison — personal biography - Infra Lead
Song Han
3 ptsMIT, model compression and quantization, efficient inference
Co-authored Deep Compression, a method for storing neural networks with fewer bits by pruning connections and quantizing their weights. The team's experiments compressed image recognition networks while preserving measured accuracy, testing whether trained models could fit into less memory.
Deep Compression - Infra Lead
Soumith Chintala
3 ptsPyTorch co-creator
Co-created PyTorch, giving researchers a Python interface for building neural networks with automatic differentiation and accelerated computation. Co-authored the PyTorch library paper, explaining the design choices behind its imperative programming style and how the implementation supports both experimentation and performance.
PyTorch: An Imperative Style, High-Performance Deep Learning Library - Safety Lead
Stuart Russell
2 ptsWrote THE AI textbook, human-compatible AI advocate
Co-authored cooperative inverse reinforcement learning, a framework for a person and a robot learning to work toward the person's objectives. The robot begins uncertain about those objectives. Human actions, including teaching, give it information about what it should try to achieve.
Cooperative Inverse Reinforcement Learning - Evals Lead
Summer Yue
2 ptsFrontier model evaluation and red-teaming research
Co-authored MultiChallenge, a benchmark that tests how well models follow a conversation over multiple turns. The tasks spread requirements through a dialogue, checking whether a model can keep track of earlier instructions and relevant context rather than treating each new message in isolation.
MultiChallenge - Infra Lead
Susan Zhang
2 ptsLed the OPT-175B training run
Co-authored OPT, a family of language models released for research. The paper credits her with initial planning and work on training infrastructure and monitoring OPT-175B. The team released model code alongside a training logbook. That record documents infrastructure problems encountered during training, giving researchers more context than a finished model alone.
OPT: Open Pre-trained Transformer Language Models - Wildcard only
swyx (Shawn Wang)
1 ptLatent Space, named the AI Engineer
Wrote The Rise of the AI Engineer and Learn in Public, about building with AI and sharing what you learn. Through Latent Space and AI Engineer events, he documents how developers turn model capabilities into working software, with an emphasis on the people doing that work.
swyx — work and writing - Post-training Lead
Teknium
2 ptsCo-founded Nous Research, open finetuning at community scale
Co-authored the Hermes 3 technical report under the name Ryan Teknium, documenting the team's work on instruction-tuned language models. The report describes training for conversation and tool use, with public benchmark results that let readers examine the behavior of the released models.
Hermes 3 Technical Report - Infra Lead
Tianqi Chen
4 ptsCreated XGBoost and TVM, built MLC for on-device inference
Developed XGBoost with Carlos Guestrin, combining boosted decision trees with systems work on sparse data, memory access and distributed processing. Co-authored TVM, an open compiler that optimizes neural-network computations for different hardware, including CPUs, GPUs and specialized accelerators.
XGBoost: A Scalable Tree Boosting SystemTVM: An Automated End-to-End Optimizing Compiler for Deep Learning - Infra Lead
Tim Dettmers
2 ptsCreated QLoRA, bitsandbytes, and 8-bit optimizers
Co-authored QLoRA, a method for tuning large language models while keeping most model weights frozen in a compact numerical format. Created bitsandbytes, a library of operations that reduce memory use. Together with QLoRA, this work addresses fitting large models onto the hardware available for training and inference.
QLoRA: Efficient Finetuning of Quantized LLMsbitsandbytes — project repository - Chief Scientist
Tom Brown
3 ptsFirst author of the GPT-3 paper, Anthropic co-founder
Was first author of the GPT-3 paper, testing whether a language model could perform new tasks from examples in its prompt without further weight updates. Co-authored Deep Reinforcement Learning from Human Preferences, using people's comparisons of behavior clips to help specify what an agent should learn.
Language Models are Few-Shot LearnersDeep reinforcement learning from human preferences - Research Lead
Tri Dao
4 ptsFlashAttention and Mamba co-creator
Co-created FlashAttention, reducing transfers between GPU memory levels while still computing exact attention. The work targets a bottleneck in training on long sequences. With Albert Gu, wrote Mamba, designing selective state-space models that retain or forget information according to the input.
FlashAttentionMamba: Linear-Time Sequence Modeling with Selective State Spaces - Safety Lead
Victoria Krakovna
2 ptsLed research on specification gaming in AI systems
Co-authored DeepMind's explanation and collection of specification-gaming examples, where agents find ways to satisfy a reward while missing the intended task. The examples make the failure concrete: a formally successful behavior can still differ from the outcome the designer wanted.
Specification gaming — Google DeepMind - Wildcard only
Wang Xiaochuan
2 ptsFounded Sogou, then founded Baichuan
Co-authored the Baichuan 2 report, where he is listed as Xiaochuan Wang. The report describes multilingual base and chat models, how they were trained, and evaluations on language tasks and specialized subjects. It provides technical context for the models released under the Baichuan name.
Baichuan 2: Open Large-scale Language Models - Evals Lead
Wei-Lin Chiang
2 ptsCo-created Vicuna and Chatbot Arena, co-founded LMArena
Co-authored Chatbot Arena, which asks people to compare responses from pairs of models. The paper examines how to turn those preferences into model comparisons, including the statistical methods and limits of crowdsourced judgments. The resulting rankings describe the collected preferences, not every aspect of model quality.
Chatbot Arena - Infra Lead
Woosuk Kwon
2 ptsCo-created vLLM, first author of the PagedAttention paper
Co-authored PagedAttention and co-created vLLM, work aimed at using GPU memory more efficiently when serving language models. PagedAttention stores attention history in blocks. The design reduces wasted memory so a serving system can handle more requests with the same hardware.
Efficient Memory Management for Large Language Model Serving with PagedAttentionIntroducing vLLM — original project announcement - Wildcard only
Yan Junjie
2 ptsFounded MiniMax, computer vision research background
Co-authored a paper on detecting faces from different viewing angles with aggregate channel features. The method combines visual signals such as image gradients and tests how efficiently they can locate faces, connecting his computer-vision research with a specific detection problem.
Aggregate channel features for multi-view face detection - Infra Lead
Yangqing Jia
3 ptsCreated Caffe, co-created ONNX, founded Lepton AI
Created Caffe and co-authored its paper, giving researchers a framework for training convolutional networks and putting them into use. Caffe combined a C++ core with Python and MATLAB interfaces. Its reference models gave other researchers a starting point for reproducing experiments.
Caffe: Convolutional Architecture for Fast Feature EmbeddingCaffe — project creators and framework - Chief Scientist
Yann LeCun
5 ptsConvnets pioneer, 2018 Turing Award, JEPA world-models research
Developed convolutional networks with collaborators for recognizing handwritten characters. His early work applied learned visual features to reading tasks instead of relying entirely on hand-designed features. Co-wrote the 2015 Deep Learning review with Yoshua Bengio and Geoffrey Hinton, explaining how these methods work.
Deep learning — NatureYann LeCun: convolutional networks for character recognition - Chief Scientist
Yejin Choi
2 ptsCommonsense reasoning pioneer, MacArthur fellow
Co-authored ATOMIC, building a commonsense resource around everyday events and the intentions, reactions and consequences people associate with them. With the ATOMIC team, turned these relationships into a dataset for studying inferences such as why someone acted and how others might feel afterward.
ATOMIC: An Atlas of Machine Commonsense for If-Then Reasoning - Chief Scientist
Yoshua Bengio
4 ptsTuring Award, Mila founder, pivoted hard into AI safety
Co-authored A Neural Probabilistic Language Model, which learned word representations together with a model for predicting the next word in a sequence. Wrote Deep Learning of Representations: Looking Forward, setting out research questions about how models learn abstractions and transfer them between tasks.
Deep Learning of Representations: Looking ForwardA Neural Probabilistic Language Model - Post-training Lead
Zhihong Shao
2 ptsFirst author of DeepSeekMath, which introduced the GRPO algorithm
Was first author of DeepSeekMath, combining math-focused pretraining, instruction tuning and reinforcement learning in a language model. The team introduced Group Relative Policy Optimization, or GRPO, as part of that work and published experiments examining its use for mathematical reasoning.
DeepSeekMath - Research Lead
Zhilin Yang
3 ptsTransformer-XL and XLNet co-author, founded Moonshot AI (Kimi)
Was first author of XLNet, developing a pretraining method with collaborators that learns from different word-prediction orders to use context on both sides. Co-authored GLM, exploring a different training objective in which a language model generates missing spans of text.
GLM: General Language Model Pretraining with Autoregressive Blank InfillingXLNet: Generalized Autoregressive Pretraining for Language Understanding
Your budget. Your team.
Read the cards, choose your people, and see the lab you make.