Research
Conference Proceedings
Do Large Language Models (LLMs) Understand Chronology?
AAAI 2026 · Oral (Student Abstract) · Poster (AI4TS Workshop)
Abstract
Large language models have shown great potential as forecasting tools in finance and economics, but backtesting performance is subject to look-ahead bias if the period overlaps with an LLM’s training window. Prompt-based attempts to avoid look-ahead bias require that LLMs understand chronology. We test LLMs’ ability to understand and enforce chronological order in three types of tasks: sorting randomly shuffled historical events; conditional sorting of events defined by some conditions; and anachronism detection based on intersections of multiple timelines. Our experiments use events that we first confirm are known to the LLM; this ensures that we test chronological understanding on an LLM’s pretrained internal knowledge. Across three LLM families— GPT-4.1 (standard), GPT-5 (hybrid-reasoning), and Claude 3.7 Sonnet (large-reasoning, with and without Extended Thinking), we find that performance degrades rapidly with problem complexity but improves greatly for reasoning models with test-time extended reasoning. These patterns are important for the real-time application of LLMs in finance.
Presentations & recognition
- Oral at AAAI-26, Student Abstract & Poster Program (Top 11%).
- Poster at AI4TS: AI for Time Series Analysis (AAAI-26 Workshop).
- Invited presentation at Yale Undergraduate Research Conference (YURC 2026), IISE Annual Conference 2026, 2026 Berkeley IEOR Annual Advisory Board Meeting
- Featured as foundational literature in OpenAI’s “Scaling Social Science Research” (2026) paper - GPT as a measurement tool.
Working Papers & Preprints
CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks
NBER Working Paper #35663, 2026 · Accepted at NeurIPS 2026 AABA4ET Workshop
Abstract
Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of another (weaker) agent. We introduce a unified framework that evaluates the capability of models to automate and augment another agent’s performance. Across seven economically grounded real-world tasks, an assistant model writes assistance text for a standardized lower-capacity worker model, which produces the deliverable. In automation mode, the assistant produces the output directly. Outputs are scored through blind pairwise comparisons by an LLM judge panel with task-specific rubrics, replicated across ten runs. Rankings across the two regimes are only modestly correlated, and the automation winner loses augmentation on five of seven tasks. Assistance is not reliably positive. The unaided worker outranks every assisted condition on three tasks, and only one model’s guidance beats no guidance on average. These results suggest that automation ability is an incomplete proxy for assistance quality, motivating benchmarks that evaluate models according to the roles they play in human-AI and multi-agent systems.
Conference presentations
- Accepted at the NeurIPS 2026 Workshop on Agentic AI Benchmarks and Applications for Enterprise Tasks (AABA4ET).
- Presented at the 2026 Wharton Generative AI & Business Conference. [slides]
Contributions as an RA
Calyber: A Ridesharing Game.
INFORMS Transactions on Education
Abstract
This case introduces Calyber, a simulation-based game designed to provide a hands-on and engaging experience in developing real-time pricing and matching decisions for shared ride services, where multiple riders are pooled into a single vehicle. Students design and implement dynamic pricing and matching policies using a rich historical ridesharing data set, competing for top performance on a holdout test set. Through this case, students gain practical insight into stochastic dynamic decision making within a modern, relevant, and data-driven context. Results from previous class implementations provide strong evidence of enhanced learning and engagement.
Notes & recognition
- 🏆 Runner-up, 2025 INFORMS Case Competition
On-Off Systems with Strategic Customers
Under revision
Abstract
Motivated by applications such as rapid transit and make-to-order systems, we study a fluid model of a many-server on-off system with multiple queues. Each server visits the queues in a prescribed order. A queue is “on” when it is being served and “off” otherwise. Customers of each queue arrive over time and decide whether to join. Services are exhaustive: servers must clear each queue before departing, while the planner determines how long they remain after each queue is first cleared. We characterize the customer equilibrium as the solution to a bounded linear complementarity problem and develop an algorithm with closed-form updates that computes the unique equilibrium in at most n steps, where n is the number of queues. We then develop an O(n^2)-time algorithm, independent of the number of servers, for computing the optimal service policy. We further show that, under an optimal policy, service is extended beyond first clearance at no more than m queues, where m is the number of servers. This structural result yields a sharper optimization procedure that requires at most 2n steps in the single-server case. We also extend our analysis to settings beyond exhaustive services. We illustrate these results using data from Bay Area Rapid Transit, where we optimize train schedules and station dwell times, recovering almost all lost ridership due to prolonged waiting times without adding vehicles.
Other Research & Awards
Data-Driven Evaluation of Board of Directors Effectiveness: Unsupervised Learning and Predictive Modeling of “Skills Matrices”
Berkeley CDSS Data Discovery Symposium, 2025
Notes & recognition
- 🏆 Best Data Visualization Award at the 2025 Berkeley CDSS Data Discovery Symposium
Mixed-Integer Linear Program for Options Pricing and Portfolio Optimization
Wells Fargo & Berkeley IEOR Bay Area Decision Sciences Summit, 2025
Notes & recognition
- 🏆 1st Runner-Up at the 2025 Wells Fargo & Berkeley IEOR Bay Area Decision Sciences Summit
- Also presented at the 2025 Berkeley IEOR Community Celebration & Alumni Achievement Ceremony
Works in Progress
Flex or Fast? Incentive-Compatible Demand Allocation in Destination-Mode Ride-Hailing Networks
Please refer to my CV for more detailed and complete research assistantships & publications.
Research interests
Data-Driven Service Operations & Market Design
I study the design of dynamic service systems in which information is asymmetric and participants respond strategically to prices, incentives, and system conditions. I combine tools from stochastic optimization, game theory, and empirical methods to model and analyze operational decisions, focusing primarily on public-sector problems such as urban transportation, as well as pricing and matching in online platforms.
On the application side, I developed dwell-time allocation policies for San Francisco Bay Area Rapid Transit (BART) that helped improve throughput on its Yellow Line while accounting for strategic passenger arrivals. I also helped design Calyber: A Ridesharing Game (2025 INFORMS Case Competition Runner-up), a case study deployed in a graduate-level supply chain course at Berkeley, where students develop dynamic pricing and matching policies for a Chicago ride-hailing company.
I believe research can and should extend beyond theory. I am particularly interested in designing and improving routing and matching systems across logistics, transportation, marketplaces, and exchanges because they shape how people, goods, and resources are allocated. Looking ahead, I hope to collaborate with practitioners and policymakers to translate my research into socially impactful solutions to real-world challenges.
AI for Operations
I investigate the reliability of generative AI in high-stakes decision systems. My recent work (AAAI 2026) empirically audits the limits of LLMs in chronological reasoning, with implications for mitigating lookahead bias in forecasting tasks.
As AI capabilities advance, I believe interdisciplinary research on how AI augments human judgment will become increasingly important. Models can be seen solving complex, well-specified problems, but they cannot yet determine which questions are worth asking, which assumptions matter, or how technical decisions will affect people. In my recent talk (Research Perspectives on the Capabilities, Limits, and Future of AI), I explore these questions and their future implications for the field.
Operations for AI
I study how firms should optimally design workflows and allocate tasks among humans, autonomous models, and AI assistants given differences in their capabilities, costs, speeds, and reliability. In CentaurBench, we show that a model's ability to automate a task is distinct from its ability to assist another agent. Some frontier models excel at automation but perform poorly as assistants, highlighting the need to benchmark models for the roles they play within a workflow, not just their standalone performance.
More broadly, I view integrating intelligence into enterprise workflows as a multifaceted operations problem, and not just a model selection problem. Tasks arrive dynamically and must be matched to heterogeneous agents whose performance may vary with workload, context, and time. Through the lenses of operations management and research, I hope to formalize and optimize these complex, evolving systems.
Professional Experience
Please refer to my resume for more recent industry experiences.