Manya Wadhwa

Manya Wadhwa

I am a PhD student at NYU (as a part of the TAUR lab), co-advised by Professor Greg Durrett and Professor Jessy Li. I am currently a Visiting Research Scientist at Meta/FAIR in NYC, working with Jesse Dodge. I am currently supported by the AIM Fellowship (AI Fellowship by Meta).

My interests lie in evaluating language models (CREATE, GENIE, EvalAgent, SmoothScore) and in finding ways to train and improve them (Detect-Critique-Refine, SkillFactory, Set-level RL). My focus has primarily been on non-verifiable tasks like short-form QA, long-form writing, scientific ideation, and creative problem solving.

I started my PhD at UT Austin, where I was a part of the UT NLP group.

Experience

During my PhD, I have worked with Adobe (Document Intelligence team) and Salesforce Research (Interactive AI team).

Before my PhD, I was an Associate Research Engineer with the R&D team at Goldman Sachs for three years. This team was led by Vijay Saraswat. I closely worked with Johannes Hoffart and Luciano Del Corro on problems in the space of Document AI and Conversation Understanding specific to the field of Finance.

I completed my Master's in Computer Science from Johns Hopkins University where I worked with Professor Mark Dredze. I did my undergrad from IIIT-Delhi where I majored in Computer Science.

News

  • Started as a Visiting Research Scientist at Meta/FAIR in NYC!
  • Interning with Salesforce Research in CA!
  • Co-presenting SkillFactory at ICLR 2026! See you in Brazil!
  • I am grateful to have also been selected for the Bloomberg Data Science PhD Fellowship (could not accept due to a prior commitment).
  • Moved to NYU as a fourth year PhD student with my advisor!
  • Cleared my qualification presentation at UT Austin!
Older news
  • EvalAgent and QUDSim accepted to COLM 2025! See you in Montreal! 🇨🇦
  • Learning to Refine accepted to EMNLP 2024 Findings! See you in Miami! 🌴🦩
  • First PhD paper (Explanation-Based Rescaling) accepted to COLM 2024! See you in Philly!
  • Started Internship at Adobe Research, with the Document Intelligence Team!

Publications

Training Reasoning Models to Generate Diverse Scientific Ideas
Under review
Manya Wadhwa, Pranav Narayanan Venkit, Hiroaki Hayashi, Yu Li, Tuhin Chakrabarty, Chien-Sheng Wu
Quantifying Claim Smoothing by LLM Research Agents
Under review
Tiasa Singha Roy, Manya Wadhwa, Greg Durrett
GENIE: A Fine-Grained Measure for Novelty
EMNLP 2026 (to appear)
Ramya Namuduri, Manya Wadhwa, Anshun Asher Zheng, Greg Durrett, and Junyi Jessy Li
CREATE: Testing LLMs for Associative Creativity
Arxiv 2026
Manya Wadhwa, Tiasa Singha Roy, Harvey Lederman, Junyi Jessy Li, Greg Durrett
SkillFactory: Self-Distillation For Learning Cognitive Behaviors
ICLR 2026
Zayne Sprague, Jack Lu, Manya Wadhwa, Sedrick Keh, Mengye Ren, Greg Durrett
ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
NeurIPS Datasets and Benchmarks Track 2025
Liyan Tang, Grace Kim, Xinyu Zhao, Thom Lake, Wenxuan Ding, Fangcong Yin, Prasann Singhal, Manya Wadhwa, Zeyu Leo Liu, Zayne Sprague, Ramya Namuduri, Bodun Hu, Juan Diego Rodriguez, Puyuan Peng, Greg Durrett
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
Proceedings of COLM 2025
Manya Wadhwa, Zayne Sprague, Chaitanya Malaviya, Philippe Laban, Junyi Jessy Li, Greg Durrett
Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation
Proceedings of COLM 2025
Tuhina Tripathi, Manya Wadhwa, Greg Durrett, Scott Niekum
QUDSim: Quantifying Discourse Similarities in LLM-Generated Text
Proceedings of COLM 2025
Ramya Namuduri, Yating Wu, Anshun Asher Zheng, Manya Wadhwa, Greg Durrett, Junyi Jessy Li
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
Proceedings of ICLR 2025
Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, Greg Durrett
Learning to Refine with Fine-Grained Natural Language Feedback
Findings of EMNLP 2024
Manya Wadhwa, Xinyu Zhao, Junyi Jessy Li, Greg Durrett
Using Natural Language Explanations to Rescale Human Judgments
Proceedings of COLM 2024
Manya Wadhwa, Jifan Chen, Junyi Jessy Li, Greg Durrett
Aligning Public Feedback to Requests for Comments on Regulations.gov
ICWSM 2020
Manya Wadhwa, Silvio Amir, Mark Dredze