Trapit Bansal is an artificial intelligence researcher whose work focuses on reinforcement learning, reasoning, natural language processing, deep learning, and meta-learning. He is a founding member of TBD Lab at Meta, where he works on AI research. Bansal previously worked at OpenAI and was a foundational contributor to the company's o1 reasoning model. [3] [6] [7]
Bansal studied at the Indian Institute of Technology Kanpur (IIT Kanpur) from 2007 to 2012, completing an integrated master's degree in Mathematics and Statistics. [2] [3] [4]
He later attended the University of Massachusetts Amherst, where he earned an M.S. in Computer Science and a Ph.D. in Computer Science. He studied under Andrew McCallum, and his graduate research focused on natural language processing, deep learning, and meta-learning. [3] [4] [9]
Bansal began his professional career at Accenture Management Consulting in Gurugram, India, where he worked as an analyst from 2012 to 2013. He subsequently worked at the Indian Institute of Science (IISc) in Bengaluru, conducting research on Bayesian modelling and inference.
During his graduate studies, Bansal completed research internships at several technology companies. In 2016, he interned with Facebook's Applied Machine Learning group, working on deep learning methods for natural language processing. He interned at OpenAI in 2017, where his research focused on multi-agent reinforcement learning and meta-learning, and at Google Research in 2018, where he worked on knowledge-graph reasoning. In 2020, he interned at Microsoft Research Montréal and researched self-supervised meta-learning for natural language processing. [4] [9]
Bansal joined OpenAI as a Member of Technical Staff in January 2022. His work focused substantially on reinforcement learning and AI reasoning. [3] [5]
Bansal was involved in the research that led to OpenAI's o1 family of reasoning models. OpenAI officially lists him among the foundational contributors to o1's reasoning research. Bansal has described himself as a co-creator of OpenAI's o-series models. [6] [7]
Before the release of o1, Bansal worked on reinforcement-learning approaches to reasoning alongside researchers including Ilya Sutskever. TechCrunch reported that he was an important participant in the early reinforcement-learning research that contributed to OpenAI's reasoning-model program. [5]
Bansal had previously worked at OpenAI as a research intern in 2017. During that period, he co-authored research on competitive multi-agent reinforcement learning and self-play. The resulting paper, Emergent Complexity via Multi-Agent Competition, examined how agents trained through competition could develop increasingly complex behaviors in simulated environments. [4] [10]
Bansal left OpenAI in June 2025 after approximately three and a half years as a full-time researcher. [3] [5]
Bansal joined Meta in July 2025 as the company expanded its efforts in advanced AI and superintelligence research. [1] [3] [5]
At Meta, Bansal became a founding member of TBD Lab, a research group working on next-generation AI models. His current professional profiles describe his role as AI research at TBD Lab within Meta. [3] [7]
His move came during a period in which Meta was recruiting researchers from organizations including OpenAI, Google DeepMind, Anthropic, and other AI laboratories to strengthen its work on reasoning and advanced foundation models. [5]
In 2026, Bansal publicly discussed Meta's work on evaluating model reasoning capabilities through international academic competitions. In August 2026, he said that Meta models had been entered into five international STEM Olympiads as an evaluation of their reasoning abilities and had achieved results corresponding to gold medals in all five. According to Bansal, three of the evaluations involved live participation and official grading. [8]
Earlier, while discussing the 2026 Asian Physics Olympiad, Bansal described the competition as an evaluation of a Meta model's multimodal and reasoning capabilities. [8]
Bansal's research has covered natural language processing, reinforcement learning, meta-learning, knowledge representation, and reasoning systems.
One of his early OpenAI projects was competitive self-play. In the 2017 paper Emergent Complexity via Multi-Agent Competition, Bansal and collaborators including Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch demonstrated how competitive environments could provide an automatic curriculum for reinforcement-learning agents and result in increasingly complex learned behaviors. [10]
Bansal also worked on relational reasoning for natural language processing. In RelNet: End-to-End Modeling of Entities & Relations, he and his co-authors introduced a memory-augmented neural-network architecture designed to model entities and relationships within text. [4]
His doctoral research included work on self-supervised meta-learning for few-shot natural language processing. In research published at EMNLP 2020, Bansal and collaborators proposed generating meta-learning tasks from unlabeled text to improve generalization when only small amounts of labeled data are available. [11]
His Ph.D. dissertation, Few-Shot Natural Language Processing by Meta-Learning Without Labeled Data, focused on methods for improving the ability of language models to adapt to new tasks with limited labeled examples. [4]
The precise terms of Bansal's compensation have not been publicly confirmed. Reports have variously described the figure as a salary, signing bonus, or broader compensation package. In August 2026, The Times of India reiterated the reported ₹800 crore figure but explicitly stated that the compensation figures in its article had not been independently verified. [9]
Last updated:
On August 17, 2026. 02:24 UTC
Edit summary:
Updated wiki: removed Superintelligence Labs Jul 2025, added Superintellig .., tags[AI,Developers]->[Developer], category->People in crypto, refs3->11