About
I’m an an assistant professor of Technology, Operations, and Statistics in the Technology Group at the NYU Stern School of Business. I did my Ph.D. in Statistics at Harvard University and was a Stone Ph.D.Scholar.
I develop statistical and computational methods to improve organizational decisions and public policy and to make AI systems safer and more trustworthy. My research spans responsible AI, algorithmic decision-making, and computational social science, building conceptual and practical tools to address issues in areas like education, hiring, criminal justice, elections, and news media.
Before coming to Harvard, I was a Ph.D. student in Computational and Mathematical Engineering at Stanford University and a Knight-Hennessy Scholar. I also spent two years as a data scientist at the Stanford Computational Policy Lab, working on projects with the public and private sector partners. I hold an A.B. in mathematics from Harvard, where I studied set theory and ergodic theory, and an M.Sc. in the history of science from the University of Oxford.
Working Papers
-
Johann D. Gaebler, Calvin Isley, Christopher Avery, and Sharad Goel. “Reassessing the Role of Standardized Tests in University Admissions.” Submitted, 2026.
There is a long-running debate over using standardized test scores to inform college and graduate admissions decisions, with some arguing that test scores are an important signal of academic strength and others arguing that they are biased and exclusionary. Here we revisit this issue by analyzing a novel dataset of more than 13,000 applications over roughly a decade to a large public policy master’s program in the United States. Consistent with past work, we find that GRE scores substantially improve predictions of first-year grades relative to predictions based on GPA alone. However, when these predictions are used to inform admissions decisions, we find that test scores only modestly improve the expected academic quality of admitted students. The gap shrinks further when we augment the test-aware and test-blind predictive models with more fine-grained information available in student transcripts and other application materials. Specifically, we estimate that incorporating standardized test scores in our setting would result in admitting students who perform, on average, only 0.03 grade points better. We show—both empirically and theoretically—that this pattern stems from a subtle distinction between predictions and decisions. Even with improved predictions, the downstream admissions decisions are often the same; and where there are differences, they often involve selecting between similarly qualified applicants. Our results indicate that standardized test scores may be less important for university admissions than previously suggested.
-
Calvin Isley, Johann D. Gaebler, and Sharad Goel. “Mitigating Label Bias with Interpretable Rubric Embeddings.” Submitted, 2026.
Statistical decision algorithms are increasingly deployed in domains where ground-truth labels are hard to obtain, such as hiring, university admissions, and content moderation. In these settings, models are typically trained on historical human evaluations—for example, using past hiring decisions as a proxy for true applicant quality. However, if past evaluations unjustly favor certain groups, models trained on these labels may inherit those biases. To address this problem, we propose basing predictions on rubric embeddings, a representation framework that replaces standard black-box embeddings with features derived from expert-defined criteria that align with the underlying construct of interest. By anchoring predictions to semantically meaningful dimensions, this approach guards against biased proxy signals. We provide both theoretical and empirical evidence that rubric embeddings mitigate label bias under plausible conditions. Empirically, we evaluate our method on a novel dataset of applications to a large master’s program. We find that models trained on rubric embeddings reduce group disparities while improving measures of cohort quality. Our results suggest that basing predictions on interpretable, domain-grounded representations offers a practical approach to learning in the presence of biased labels.
Publications
-
Jongbin Jung, Sam Corbett-Davies, Johann D. Gaebler, Ravi Shroff, and Sharad Goel. “Measuring Disparate Impact in Human and Machine Decisions.” Proceedings of the National Academy of Sciences (2026).
Empirical analyses have grown increasingly important in discrimination litigation with the greater availability of detailed data on individuals and decisions. A popular analytic strategy is to estimate disparities after adjusting for observed covariates, typically with a regression model, in hopes of ferreting out discriminatory intent. This approach, however, is ill-suited to auditing algorithms that are now commonly used to aid decisions, which typically do not include race or other legally protected factors as inputs. Motivated by legal understandings of disparate impact, we introduce a new approach, which aims to measure “unjustified” disparities in both human and machine decisions. Our method, which we call risk-adjusted regression, proceeds in three steps. In the first step, we combine all available information in a machine learning model to estimate the value, or inversely, the risk, of taking a certain action, such as approving a loan application or hiring a job candidate. Second, we measure disparities in decisions after adjusting for these risk estimates alone. Finally, in the third step, we assess the sensitivity of results to potential mismeasurement of risk. We demonstrate this approach on a detailed dataset of 2.2 million police stops of pedestrians in New York City, and show that traditional statistical tests of discrimination can substantially understate the magnitude of (risk-adjusted) racial disparities.
-
Johann D. Gaebler and Sharad Goel. “A Simple, Statistically Robust Test of Discrimination.” Proceedings of the National Academy of Sciences (2025). (EAAMO 2024 Best Paper Award.)
In observational studies of discrimination, the most common statistical approaches consider either the rate at which decisions are made (benchmark tests) or the success rate of those decisions (outcome tests). Both tests, however, have well-known statistical limitations, sometimes suggesting discrimination even when there is none. Despite the fallibility of the benchmark and outcome tests individually, here we prove a surprisingly strong statistical guarantee: under a common non-parametric assumption, at least one of the two tests must be correct; consequently, when both tests agree, they are guaranteed to yield correct conclusions. We present empirical evidence that the underlying assumption holds approximately in several important domains, including lending, education, and criminal justice—and that our hybrid test is robust to the moderate violations of the assumption that we observe in practice. Applying this approach to 2.8 million police stops across California, we find evidence of widespread racial discrimination.
-
Johann D. Gaebler, Sean J. Westwood, Shanto Iyengar, and Sharad Goel. “No News is Good News? The Declining Information Value of Broadcast News in America .” PLOS ONE (2025).
Despite the rise of digital media, Americans are five times more likely to consume news via television than through online platforms. However, due in large part to technical hurdles, it remains unclear what content appears on broadcast news and how the mixture of content has changed over time. We consider these questions by applying a novel LLM-based approach to an understudied corpus of expert-generated summaries of virtually all news segments aired on the “big three” broadcast networks—ABC, CBS, and NBC—between 1968 and 2019. Results based on nearly one million news segments show that “information density”—the amount of time dedicated to political issues—has declined substantially over the last 50 years. Today, broadcast news spends twice as much time on commercials and “soft” news and half as much on issue-based political coverage compared to a few decades ago. Since the early 1990s, the news has also shifted inward, focusing more on domestic stories and less on international affairs. These changes suggest a transformation in the informative role of broadcast news, raising questions about its impact on voter knowledge and political engagement.
-
Johann D. Gaebler, Sharad Goel, Aziz Huq, and Prasanna Tambe. “Auditing large language models for race & gender disparities: Implications for artificial intelligence–based hiring .” Behavioral Science & Policy (2025).
Rapid advances in artificial intelligence (AI), including large language models (LLMs) with abilities that rival those of human experts on a wide array of tasks, are reshaping how people make important decisions. At the same time, critics worry that LLMs may inadvertently discriminate against some groups. To address these concerns, recent regulations call for auditing the LLMs used in important decisions such as hiring. But neither current regulations nor the scientific literature offers clear guidance on how to conduct these audits. In this article, we propose and investigate one approach for auditing algorithms: correspondence experiments, a widely applied tool for detecting bias in human judgments. We applied this method to a range of LLMs instructed to rate job candidates using a novel data set of job applications for K-12 teaching positions in a large American public school district. By altering the application materials to imply that candidates are members of specific demographic groups, we measured the extent to which race and gender influenced the LLMs’ ratings of the candidates’ suitability. We found moderate race and gender disparities, with the models slightly favoring women and non-White candidates. This pattern persisted across several variations in our experiment. It is unclear what might be driving these disparities, but we hypothesize that they stem from posttraining efforts, which are part of the LLM training process and intended to correct biases in these models. We conclude by discussing the limitations of correspondence experiments for auditing algorithms.
-
Sam Corbett-Davies*, Johann D. Gaebler*, Hamed Nilforoshan*, Ravi Shroff, and Sharad Goel. “The Measure and Mismeasure of Fairness.” Journal of Machine Learning Research (2023).
The field of fair machine learning aims to ensure that decisions guided by algorithms are equitable. Over the last decade, several formal, mathematical definitions of fairness have gained prominence. Here we first assemble and categorize these definitions into two broad families: (1) those that constrain the effects of decisions on disparities; and (2) those that constrain the effects of legally protected characteristics, like race and gender, on decisions. We then show, analytically and empirically, that both families of definitions typically result in strongly Pareto dominated decision policies. For example, in the case of college admissions, adhering to popular formal conceptions of fairness would simultaneously result in lower student-body diversity and a less academically prepared class, relative to what one could achieve by explicitly tailoring admissions policies to achieve desired outcomes. In this sense, requiring that these fairness definitions hold can, perversely, harm the very groups they were designed to protect. In contrast to axiomatic notions of fairness, we argue that the equitable design of algorithms requires grappling with their context-specific consequences, akin to the equitable design of policy. We conclude by listing several open challenges in fair machine learning and offering strategies to ensure algorithms are better aligned with policy goals.
-
Johann D. Gaebler, Phoebe Barghouty, Sarah Vicol, Cheryl Phillips, and Sharad Goel. “Forgotten but not gone: A multi-state analysis of modern-day debt imprisonment.” PLOS ONE (2023).
In almost every state, courts can jail those who fail to pay fines, fees, and other court debts—even those resulting from traffic or other non-criminal violations. While debtors’ prisons for private debts have been widely illegal in the United States for more than 150 years, the effect of courts aggressively pursuing unpaid fines and fees is that many Americans are nevertheless jailed for unpaid debts. However, heterogeneous, incomplete, and siloed records have made it difficult to understand the scope of debt imprisonment practices. We culled data from millions of records collected through hundreds of public records requests to county jails to produce a first-of-its-kind dataset documenting imprisonment for court debts in three U.S. states. Using these data, we present novel order-of-magnitude estimates of the prevalence of debt imprisonment, finding that between 2005 and 2018, around 38,000 residents of Texas and around 8,000 residents of Wisconsin were jailed each year for failure to pay (FTP), with the median individual spending one day in jail in both Texas and Wisconsin. Drawing on additional data on FTP warrants from Oklahoma, we also find that unpaid fines and fees leading to debt imprisonment most commonly come from traffic offenses, for which a typical Oklahoma court debtor owes around $250, or $500 if a warrant was issued for their arrest.
-
William Cai, Johann Gaebler, Justin Kaashoek, Lisa Pinals, Samuel Madden, and Sharad Goel. “Measuring Racial and Ethnic Disparities in Traffic Enforcement with Large-Scale Telematics Data.” PNAS: Nexus (2022).
Past studies have found that racial and ethnic minorities are more likely than white drivers to be pulled over by the police for alleged traffic infractions, including a combination of speeding and equipment violations. It has been difficult, though, to measure the extent to which these disparities stem from discriminatory enforcement rather than from differences in offense rates. Here, in the context of speeding enforcement, we address this challenge by leveraging a novel source of telematics data, which include second-by-second driving speed for hundreds of thousands of individuals in 10 major cities across the United States. We find that time spent speeding is approximately uncorrelated with neighborhood demographics, yet, in several cities, officers focused speeding enforcement in small, demographically non-representative areas. In some cities, speeding enforcement was concentrated in predominantly non-white neighborhoods, while, in others, enforcement was concentrated in predominantly white neighborhoods. Averaging across the 10 cities we examined, and adjusting for observed speeding behavior, we find that speeding enforcement was moderately more concentrated in non-white neighborhoods. Our results show that current enforcement practices can lead to inequities across race and ethnicity.
-
Johann D. Gaebler, William Cai, Guillaume Basse, Ravi Shroff, Sharad Goel, and Jennifer Hill. “A Causal Framework for Observational Studies of Discrimination.” Statistics and Public Policy (2022).
In studies of discrimination, researchers often seek to estimate a causal effect of race or gender on outcomes. For example, in the criminal justice context, one might ask whether arrested individuals would have been subsequently charged or convicted had they been a different race. It has long been known that such counterfactual questions face measurement challenges related to omitted-variable bias, and conceptual challenges related to the definition of causal estimands for largely immutable characteristics. Another concern, which has been the subject of recent debates, is post-treatment bias: many studies of discrimination condition on apparently intermediate outcomes, like being arrested, that themselves may be the product of discrimination, potentially corrupting statistical estimates. There is, however, reason to be optimistic. By carefully defining the estimand—and by considering the precise timing of events—we show that a primary causal quantity of interest in discrimination studies can be estimated under an ignorability condition that is plausible in many observational settings. We illustrate these ideas by analyzing both simulated data and the charging decisions of a prosecutor’s office in a large county in the United States.
-
Sabina Tomkins, Keniel Yao, Johann Gaebler, Tobias Konitzer, David Rothschild, Marc Meredith, and Sharad Goel. “Blocks as Geographic Discontinuities: The Effect of Polling-Place Assignment on Voting .” Political Analysis (2022).
A potential voter must incur a number of costs in order to successfully cast an in-person ballot, including the costs associated with identifying and traveling to a polling place. In order to investigate how these costs affect voting behavior, we introduce two quasi-experimental designs that can be used to study how the political participation of registered voters is affected by differences in the relative distance that registrants must travel to their assigned Election Day polling place and whether their polling place remains at the same location as in a previous election. Our designs make comparisons of registrants who live on the same residential block, but are assigned to vote at different polling places. We find that living farther from a polling place and being assigned to a new polling place reduce in-person Election Day voting, but that registrants largely offset for this by casting more early in-person and mail ballots.
-
Hamed Nilforoshan*, Johann Gaebler*, Ravi Shroff, and Sharad Goel. “Causal Conceptions of Fairness and Their Consequences.” International Conference on Machine Learning (2022). (ICML 2022 Outstanding Paper Award.)
Recent work highlights the role of causality in designing equitable decision-making algorithms. It is not immediately clear, however, how existing causal conceptions of fairness relate to one another, or what the consequences are of using these definitions as design principles. Here, we first assemble and categorize popular causal definitions of algorithmic fairness into two broad families: (1) those that constrain the effects of decisions on counterfactual disparities; and (2) those that constrain the effects of legally protected characteristics, like race and gender, on decisions. We then show, analytically and empirically, that both families of definitions almost always—in a measure theoretic sense—result in strongly Pareto dominated decision policies, meaning there is an alternative, unconstrained policy favored by every stakeholder with preferences drawn from a large, natural class. For example, in the case of college admissions decisions, policies constrained to satisfy causal fairness definitions would be disfavored by every stakeholder with neutral or positive preferences for both academic preparedness and diversity. Indeed, under a prominent definition of causal fairness, we prove the resulting policies require admitting all students with the same probability, regardless of academic qualifications or group membership. Our results highlight formal limitations and potential adverse consequences of common mathematical notions of causal fairness.
-
William Cai, Johann D. Gaebler, Nikhil Garg, and Sharad Goel. “Fair Allocation through Selective Information Acquisition.” Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (2020).
Public and private institutions must often allocate scarce resources under uncertainty. Banks, for example, extend credit to loan applicants based in part on their estimated likelihood of repaying a loan. But when the quality of information differs across candidates (e.g., if some applicants lack traditional credit histories), common lending strategies can lead to disparities across groups. Here we consider a setting in which decision makers—before allocating resources—can choose to spend some of their limited budget further screening select individuals. We present a computationally efficient algorithm for deciding whom to screen that maximizes a standard measure of social welfare. Intuitively, decision makers should screen candidates on the margin, for whom the additional information could plausibly alter the allocation. We formalize this idea by showing the problem can be reduced to solving a series of linear programs. Both on synthetic and real-world datasets, this strategy improves utility, illustrating the value of targeted information acquisition in such decisions. Further, when there is social value for distributing resources to groups for whom we have a priori poor information—like those without credit scores—our approach can substantially improve the allocation of limited assets.
-
Johann Gaebler, Alexander Kastner, Cesar Silva, Xiaoyu Xu, and Zirui Zhou. “Partially Bounded Transformations have Trivial Centralizers.” Proceedings of the American Mathematical Society (2018).
We prove that for infinite rank-one transformations satisfying a property called “partial boundedness,” the only commuting transformations are powers of the original transformation. This shows that a large class of infinite measure-preserving rank-one transformations with bounded cuts have trivial centralizers. We also characterize when partially bounded transformations are isomorphic to their inverse.
Other Manuscripts
-
Johann D. Gaebler. “The Statistics of Discrimination.” Ph.D. thesis, Harvard University, 2026.
Policymakers, regulators, courts, and the public increasingly rely on statistical evidence to understand and address human and machine bias in domains like lending, policing, hiring, and medicine. Empirical studies measuring discrimination, however, face both definitional and data challenges. First, mathematical definitions of bias often imperfectly capture legal and intuitive notions of discrimination, and seemingly natural measures can have counterintuitive implications that would harm the groups they are intended to protect. Second, discrimination is often most concerning in settings where data are limited: Key information about decisions, such as the factors on which they were based, is often unavailable or difficult to measure. This dissertation develops theory, methods, and applications across a variety of domains aimed at addressing these twin challenges in the measurement of discrimination.
Part I studies algorithmic discrimination. Machine decisions are simpler to analyze than human decisions in important respects—decision inputs are structured and usually observable, and the same inputs reliably yield the same decisions—but the meaning of algorithmic “discrimination” is often unclear. Chapter 2 critiques popular mathematical definitions of algorithmic fairness that seek to equalize model performance metrics like false positive rates across race, gender, or other characteristics. Because these measures suffer from inframarginality, a subtle statistical limitation, equalizing them typically requires making suboptimal decisions, including for protected groups. Chapters 3, 4, and 5 develop alternative approaches to measuring discrimination in algorithmic decisions. In contrast to the approaches critiqued in Chapter 2, Chapter 3 advocates for a consequentialist approach to fairness, proposing an algorithm that efficiently reduces inequality while balancing competing policy goals in lending. Chapters 4 and 5 draw on legal frameworks governing discrimination. Chapter 4 develops a risk-based measure of disparate impact, along with a regression-based framework for estimating it in observational settings. Motivated by emerging AI regulations, Chapter 5 adapts correspondence experiments, a behavioral science method for detecting discrimination, to audit large language models used in hiring.
Part II extends the methods of Part I to human decision-making. Unlike algorithms, human decisions are often poorly documented, difficult to reconstruct, and costly to manipulate, hindering the measurement of discrimination. We focus on applications to criminal justice, where these problems are particularly acute. Chapter 6 develops the robust outcome test, a method for measuring disparate impact in settings where little is known about the factors driving decisions, rendering the regression-based approach of Chapter 4 infeasible. By combining “benchmark” and “outcome” tests, standard but individually flawed methods of detecting discrimination, the robust outcome test yields strong statistical guarantees under a nonparametric assumption that empirical evidence suggests often holds in practice. Paralleling the correspondence experiments of Chapter 5, Chapter 7 develops a causal framework for estimating discrimination from observational data in multi-stage decision processes like prosecutorial charging decisions occurring downstream of potentially discriminatory earlier decisions like racially biased arrests. Finally, an alternative response to data limitations is leveraging novel data sources. Chapter 8 uses large-scale telematics data, second-by-second records of driving behavior for hundreds of thousands of individuals, to benchmark racial disparities in traffic enforcement against actual differences in speeding, disentangling discriminatory enforcement from differences in underlying behavior.
-
Guanting Chen, Johann D. Gaebler, Matt Peng, Chunlin Sun, and Yinyu Ye. “An Adaptive State Aggregation Algorithm for Markov Decision Processes.” arXiv preprint, 2021.
Value iteration is a well-known method of solving Markov Decision Processes (MDPs) that is simple to implement and boasts strong theoretical convergence guarantees. However, the computational cost of value iteration quickly becomes infeasible as the size of the state space increases. Various methods have been proposed to overcome this issue for value iteration in large state and action space MDPs, often at the price, however, of generalizability and algorithmic simplicity. In this paper, we propose an intuitive algorithm for solving MDPs that reduces the cost of value iteration updates by dynamically grouping together states with similar cost-to-go values. We also prove that our algorithm converges almost surely to within \(2\varepsilon / (1 - \gamma)\) of the true optimal value in the \(\ell^\infty\) norm, where \(\gamma\) is the discount factor and aggregated states differ by at most \(\varepsilon\). Numerical experiments on a variety of simulated environments confirm the robustness of our algorithm and its ability to solve MDPs with much cheaper updates especially as the scale of the MDP problem increases.
-
Johann D. Gaebler. “Large Cardinals and Projective Determinacy.” Undergraduate thesis, Harvard University, 2017.
This thesis comes out of two streams of motivation: purely mathematical considerations, and considerations relating to the history and philosophy of mathematics. Both aspects are reflected in the two very different, but closely interconnected halves of the present work.
The mathematical goal of this thesis is to provide a thorough and accessible exposition of a series of important results which, through the axiom of determinacy, link descriptive set theory—loosely speaking, the study of well-behaved subsets of the real line—to large cardinal axioms—that is, axioms which extend the “height” of the universe of sets. This portion of the thesis, Part II, culminates in a proof that \(\mathbf{I}_0\) implies projective determinacy, which has never before appeared in print.
The non-mathematical aim of this thesis is to try to understand these results in their broader intellectual context. The historical portion of this thesis, Part I, examines the mathematical currents that coalesced into descriptive set theory in the late 19th and early 20th centuries. This historical development is illustrated through two case studies: first, the confluence of factors which led to the adoption of the infinite into mathematics as a legitimate object of study; and second, the French and Russian analysts who, at the turn of the 20th century, navigated between the revolutionary theory of sets first developed by Georg Cantor in the last decades of the 19th century and the traditions of mathematical analysis stretching back to the 17th century.
While very different in nature, it is the author’s sincere hope that it will be evident how Part I and Part II are rooted in a common collection of interests and concerns and form a unified intellectual project.
(*: Denotes equal contribution.)