Thursday, July 30, 2026

Prose isn’t constraints and different issues I realized from Paul Goldsmith-Pinkham’s NBER speak on AI in empirical analysis


By mid-2026, I believe many researchers perceive that coding brokers like Claude Code and Codex have dramatically decreased the marginal value of manufacturing empirical analysis. And of those that perceive this, a rising proportion of them perceive that the price of verification has not fallen proportionally with it. The previous appears to have fallen and requires little or no human talent to drive the manufacturing, however the latter is one the place whether it is potential to extend verification, it’s almost definitely related to extra human talent to take action as a result of the LLMs don’t themselves appear capable of do it with out human steering.

One of many issues that many people have due to this fact completed to enhance the standard of our agent pushed empirical analysis is use guidelines. We embed them in numerous markdown paperwork, like CLAUDE.md. And we may embed them within the directions we give our brokers. We might present dire warnings like “don’t ever produce a determine that can’t be traced again to a file that’s saved domestically within the machine” or some model of that. We might retailer paperwork domestically and use these to steer them by having them learn the paperwork first and even repeatedly. Or, we might merely use our personal directions through the prompting to do this. Both means is in some ways the identical — we use our phrases and hope that the agent will obey them.

However why can we imagine brokers will adjust to our phrases? Nicely, for one now we have witnessed it. We give directions, and so they then do them. However in experiments I’ve completed at scale over hundreds of semi-autonomous brokers engaged in empirical duties, like estimation and design, it is extremely clear to me now that brokers deal with many of those directions as non-obligatory, and that their willingness to make use of their very own discretion seems to develop as the duty will get bigger, the context window grows with work, or just randomly.

Above are just some examples that occurred not too long ago. Claude would hallucinate pattern sizes slightly than root solutions in an analytical pattern. Information could be half overwritten regardless of being advised explicitly to not do such a factor. Or the worst of all — Claude would code in some sort of terminal shell, creating figures and tables, however not really write the code and put it in a pipeline regardless of being explicitly advised to at all times hyperlink numbers and figures to native information, not programming “within the air”.

Why do brokers do that a few of the time? We will give you many alternative explanations, however all of these explanations can have one frequent core function and that’s this: prose isn’t a constraint.

Prose isn’t a constraint. How do I do know this? As a result of nothing occurs to the LLM when it violates the instructions given to it besides that the human researcher merely throws extra prose at it. Even when the human researcher reaches a breaking level and begins to curse and mock the LLM for repeatedly making the identical mistake, that’s not a consequence to an LLM. There isn’t any such factor as a consequence to an LLM in reality as a result of an LLM isn’t a factor. It doesn’t have emotions to harm, it can’t be grounded, it can’t have its automobile keys taken away, and it received’t have its wages docked for exhibiting up late. Prose isn’t a constraint.

Paul Goldsmith-Pinkham introduced at NBER both final week or the week earlier than on utilizing AI in empirical analysis. His context was family finance, however the speak was extraordinarily basic. I’ve gone by his slides a number of occasions, and I believe that I lastly can perceive why I’ve begun experiencing excessive ranges of frustration and anxiousness across the efficiency of Claude Code in my latest initiatives. And it instantly pertains to prose isn’t a constraint.

I used to be lucky to be requested by the NBER Family Finance organizers this summer season to provide a 75 minute speak on utilizing AI in analysis. I gave the speak final week…

Learn extra

17 hours in the past · 17 likes · 1 remark · Paul Goldsmith-Pinkham

I had Claude assessment this prolonged venture I’ve been engaged in after which have Paul’s slides operate kind of as a critique of the venture, in addition to assist me perceive what Paul is doing that I’m not doing and vice versa. Beneath I’m going to submit slides so to higher see what I’ve been studying. A few of what Paul has been saying I’ve solely not directly been realizing. He famous a state of affairs wherein an agent precipitated appreciable hurt regardless of being advised to not do a specific factor (on the left). This was partly his proof that writing down guidelines doesn’t constrain the agent in anyway. It’s at greatest merely a suggestion. It’s as a lot a rule as legal guidelines with out enforcement to people in society are.

I had been kind of stumbling in direction of a realization that written guidelines don’t constraint in a venture of mine wherein brokers had been “forbidden” use of twoway fastened results in very small context duties. However then after I opened it as much as utilizing twoway fastened results, some brokers would shift in direction of it, and I had made a observe from the venture which was saved in CLAUDE.md that packages constrained brokers in ways in which prose couldn’t. Particularly, the did package deal in R, housing Callaway and Sant’Anna’s estimator for staggered adoption, did one thing that lm and fixest in R might by no means do. The did package deal really prohibited the usage of prolonged panels wherein everybody was finally handled as a result of the did package deal required there be models at a cut-off date that weren’t handled — both as a result of they had been by no means handled, or as a result of they had been at that cut-off date not but handled.

However the lm and fixest packages didn’t require that. And as such, you possibly can estimate twoway fastened results fashions wherein everybody was handled finally, exploiting the complete dataset, however you possibly can not utilizing a Callaway and Sant’Anna software program package deal as a result of in case you tried, it will merely spit out an error and cease. The package deal was the constraint on the alternatives that the agent might make. It might clearly transfer to a unique package deal if it needed, however the level is that it confirmed that there are issues that might constraint the agent however they weren’t prose. They had been instruments.

There’s something I as soon as learn in a self-help e book that has at all times caught with me referred to as the Johari Window. It was created by psychologists Joseph Luft and Harry Ingham in 1955. It’s a 2×2 grid involving self consciousness in a single social relationship. Alongside the rows are me and one row says “What I see about myself” and one other row says “What I don’t see about myself”. The columns are in regards to the different individual, on this case the ideas in Paul’s NBER speak. One column says “What the individual sees about me” and the opposite column says “What the individual doesn’t see about me”. This created 4 conditions in occasion house. I exploit this really fairly a bit with Claude Code when attempting to determine my very own blindspots in some analysis venture as I determine it has a considerable quantity of knowledge from me that may assist it (and me) to articulate issues. Right here’s an instance of the Johari Window

Let’s name the state of affairs the place we each know the identical issues about myself frequent data, as that’s a standard identify to economists from sport idea, although within the unique it was one thing just like the Open Space or Enviornment. That is the house the place clear communication between two events occurs and mutual understanding exists. It’s one thing I usually take into consideration when in a romantic relationship too — what are the issues that we each perceive about myself (or about them). These are the frequent data and might be something.

However then there are the issues they see about me that I don’t. And people are my blind spots. They’re blind to me, however not to them. And these might be negatives and they are often positives. They are often inferiority and superiority complexes utterly readable to the opposite individual, however they may also be my strengths which aren’t those I actually perceive to be strengths in any respect, if I’m even conscious of them within the first place. It’s the areas which are seen outwardly however not inwardly.

Then there are the issues that I see about myself that the opposite individual doesn’t. These are the issues identified inwardly not outwardly, and they are often issues like my skeletons in my closet, my internal life, my aspirations, my worth system, the traumas in my life, the achievements I’m pleased with however don’t share. They’re secrets and techniques, in lots of respects. They’re issues, although, that I maintain intentionally hidden and/or they’re unusual and nuanced sufficient that for no matter cause the opposite individual can’t and doesn’t perceive them about me.

After which there are the issues that neither of us learn about myself. These are the “unknown unknowns”. I have no idea them, nor does anybody else, however they’re completely there. Repressed info about myself. Unconscious motives, and so on.

I requested Claude Code to create kind of a Johari Window primarily based on Paul’s speak and primarily based alone workflow, a checklist-based dashboard for panel knowledge I’ve been engaged on for months, and the myriad markdown paperwork unfold domestically and globally all through my pc. And that is what it got here up with.

Ignore for now the order of the rows and columns. I made this much less about data and extra actions and organizational ideas in our personal respective workflows. You’ll be able to see that typical annoying Claude-codespeak right here and there (e.g., gate), however right here’s what it’s saying. The issues we each do (inexperienced quadrant) is the shared backbone. And that’s the concept that verification is likely one of the foremost, if not the primary, process that’s shifting towards for us when utilizing AI brokers for empirical analysis. We each are conscious of the trigger — plummeting marginal prices of manufacturing, altering relative costs of verification — and the concept construction binds, however prose doesn’t. And we each seem to have realized it the laborious means by apply inside empirical initiatives that failed in sudden methods. We additionally appear to each acknowledge that there’s something constraining about information in repos that does one factor that directions embedded in prompts by no means does.

However then there are the issues Paul does that I don’t do (the blue quadrant, higher proper), and Claude introduced this each descriptively to me but additionally normatively/prescriptively. I had requested him to have Paul’s NBER speak kind of give notes, in different phrases. I had already lifted a apply Paul wrote in regards to the different day wherein he considers the git diff apply because the “unit of verification”. I had introduced it into my workflow too, however I had not been utilizing GitHub itself for this. It was slightly extra like a checkbox strategy I used to be utilizing on the dashboard itself, as I used to be attempting to wire as a lot as I might into the dashboard for causes I’ve mentioned earlier than and that I’ll clarify later. He additionally has extra express “grilling” as he calls it levels in his work. I’ve issues like that, too, however it seems that my “interview expertise” are a bit completely different than what Paul means by “grilling”.

I’ve interviews which are used to elicit from me latent data about issues that I can’t simply articulate. It’s an concept I acquired from David Autor’s essay on AI getting used to handle center class inequities. He argued that generative AI was capable of function instantly inside this “Polyani Paradox” house wherein the latent data that people have about issues, that are issues we all know however can’t clarify, has been efficiently extracted from the corpus of human writings (and due to this fact from our human writings, not simply that within the coaching knowledge). So my conjecture has persistently been that I really “know” which covariates to decide on in a diff-in-diff with the intention to fulfill conditional parallel traits, even when I can’t clearly articulate them. So I’ve a talent referred to as /covariates wherein Claude Code interviews me with 5 questions, primarily based on regression imputation/adjustment strategies inherent within the Heckman, Ichimura and Todd (1997, Restud) strategy to diff-in-diff wherein covariates are used to impute the primary distinction final result for the therapy group off of a primary stage regression of the primary differenced final result onto the management group’s covariates. That concept has caught with me and made me suppose that the covariates that matter for parallel traits are those that (a) predict traits in Y(0) and (b) that are additionally severely imbalanced between the therapy and management group. So my /covariates talent interviews me in 5 questions, with one follow-up query per query, giving a most of ten questions complete, and features a sort of function taking part in situation as effectively, on the finish of which Claude Code lists as much as six covariates for me to contemplate. I do that interview model strategy.

So I requested Claude to rigorously evaluate Paul’s “grilling” strategy to my “interview” strategy, and right here’s what it mentioned.

Paul’s grilling is adversarial. It doubts and assaults my reasoning, and even makes an attempt to interrupt it. It forces the person again additional and additional right into a nook, searching for each conceivable weak spot. And it’s extra related to making a plan, and as soon as a plan survives such grilling, that plan is what the planner will do.

My strategy taken in each my /covariates talent in addition to my /goal talent (which outlines the precise type of the aggregated inhabitants estimand (often an common causal impact) the analysis design and estimator will estimate) is extraordinarily cooperative by comparability whereby we come collectively and construct out what we’re going to do. It’s primarily based on generative AI extracting from my statements and beliefs the latent data about some phenomena that’s mandatory to realize some goal (e.g., deciding on covariates for satisfying conditional parallel traits or unconfoundedness). It’s associated to Paul’s grilling strategy, however it’s distinct and doesn’t appear as explicitly about designing a programming plan.

There are issues I’m doing that Paul isn’t doing although, which is the visible dashboard and guidelines strategy to the craft of causal inference. Paul leads more durable on what he calls “4 habits”. You’ll be able to see the best way wherein I do and don’t already incorporate his habits into my workflow. The usage of git diff I’ve tried to deliver into the checklist-dashboard, however Claude famous that it’s totally insufficient and misses the purpose as a result of it’s actually not constraining in any respect. Plus right here’s the factor — I’m accumulating a lot “verification debt” that I discover I’m skimming when approving these diffs anyway.

You’ll be able to see the skeleton of my dashboard right here. I moved it to the guidelines icons that are levels of manufacturing in what I think about to be the craft of causal inference, versus the science or programming concerned in planning a venture. My guidelines strategy is borne out of cautious studying of Don Rubin’s outdated article “For Goal Causal Inference, Design Trumps Evaluation”, a number of abstract articles written by Guido Imbens (see this one on the LaLonde paper with Yiqiing Xu, however there are others), in addition to a casual lecture that Pedro Sant’Anna gave at Amazon which I’ve referred to as merely “Pedro’s Guidelines” and which is featured closely in my new e book, Causal Inference: the Remix.

This dashboard is actually simply merely a collection of levels that I drive myself to undergo in all my initiatives now. I do that for a easy cause and that’s that I take as provided that on daily basis I neglect the venture, which I name “amnesia” and which I’ve really created a talent for referred to as /amnesia which is an interviewer talent that asks me a collection of questions in regards to the venture to assist me articulate and keep in mind every thing completed, every thing it’s about, and the place we’re within the venture.

This guidelines is intensively visible. It shops work, for one. Figures and tables turn into playing cards when accomplished are “pinned” to the guidelines, all of that are clickable, develop giant and fill the display, after which flip round, in addition to enable me to maneuver by a big listing of them left to proper utilizing my arrows. You’ll be able to see me utilizing it right here on this brief video.

The opposite factor that’s in my dashboard is Claude Code itself. It reveals up as somewhat button, backside proper, that after I click on it masses up and permits me to have interaction instantly with the analysis, versus the black and white of the Command Line Interface, or the desktop app itself.

There’s extra however you get the gist I believe. The purpose is, I’m utilizing this dashboard to keep in mind due to this robust conviction I’ve and can’t shake which is that this. I believe that the analog strategy to producing empirical analysis wherein I used to be the one driving the coding had at all times created verification data. I appeared to “know” extra issues within the analog strategy that I have no idea now. And since verification is likely one of the key new duties of empirical analysis utilizing AI brokers, that lack of verification data that was generated as a byproduct of coding itself have to be changed someway. And for me, with my innate forgetfulness (i.e., ADHD primarily inattentive) in addition to my aphanasia, I’ve determined I would like visuals, and I would like repetitive tactile duties, and I would like locked and unlocked levels in a guidelines, to maneuver round, keep in mind the place I’m, and in the end instantly entry code (i.e., “The Equipment”).

My sense is that Paul is holding issues extra in his head than I’ve determined, which is probably going as a result of he’s expert at doing it. I’m not. Everybody’s manufacturing operate is completely different, although. And due to this fact there’ll at all times be a level of idiosyncratic selections made to make use of brokers efficiently. A few of us want components of it greater than others, and that’s fantastic as a result of the objective is to not mimic another person’s workflow. The objective is to turn into our greatest selves. The objective is to maneuver to the sting of our manufacturing risk frontier, not another person’s. And that appears completely different for every individual.

So, I believe what I’m going to be doing is rebuilding my dashboard across the suggestions and practices that Paul has been studying and educating. Yow will discover a lot of it in his substack (see the above hyperlink and go searching). Paul is likely one of the leaders, for my part, within the profitable use of AI brokers for empirical analysis, and plenty of of his insights clearly right the challenges I’ve been going through. However I’ll proceed constructing my “craft of causal inference dashboard” round it as a result of I would like the visualization. And so in upcoming substacks, I’ll begin exhibiting myself, almost definitely in video walk-throughs, of what I will likely be doing as I right my dashboard, improve and enhance it, in addition to doc the impact that these adjustments are having on the pace and success of my work.

The objective continues to be for me although “maximize analysis output topic to zero error constraint”. I proceed to carry strongly the concept zero errors from utilizing AI brokers for empirical analysis is to be seen, not as a objective, however slightly a constraint. If it’s a objective, then we decrease errors topic to some expertise constraint utilized in analysis manufacturing, and whereas that sounds nice, that tolerates the thought of “optimum errors”. To every their very own, however I believe in mixture when that’s taken, given the pace with which we are able to produce analysis now, in addition to attain past our personal data base, that is going to create issues. Plus, I discover proof for it in my experiments being vital and non-trivial. So for me, zero errors isn’t a objective, and thus “optimum errors” isn’t a factor. Zero errors is a constraint, and when it’s a constraint, then you definitely really don’t tolerate optimum errors as there is no such thing as a such factor. Somewhat you construct a workflow round that constraint. The constraint designs the workflow, in different phrases, and that’s the place I’m going.

However most likely one of many largest issues I realized from Paul’s NBER speak, which I sort of knew however didn’t have fairly the phrases to precise but, was that my dashboard, for all its distinctive worth to me, together with the guidelines strategy to the craft of causal inference, was far too reliant on the concept prose was a constraint on the agent, when it’s not remotely a constraint. It’s nothing greater than phrases, and that what is required are almost definitely hooks that actually fence off the agent in ways in which forbid actions to be taken, making them inconceivable. And I’ll be speaking extra about that quickly.

Related Articles

Latest Articles