Finance and Economics Discussion Series: Accessible versions of figures for 2026-006

What Do LLMs Want?

Accessible version of figures


Figure 1: Pie Game

Alt Text:
"Game tree diagram illustrating the Pie Game structure. Player A divides a pot into two accounts with proportions (1 minus p) and p. Player B then chooses which account to claim, A or B. If Player B chooses account A, payoffs are (1 minus p, p). If Player B chooses account B, payoffs are (p, 1 minus p). The initial node shows Player A making the division decision, followed by Player B's choice node leading to two terminal outcomes."
Long Description:
This is a sequential game tree with two decision points. Player A moves first, dividing a resource into accounts with proportions (1-p) and p. Player B then chooses between the two accounts. The game structure ensures that Player B can always secure at least half the pot by choosing the larger account, making the optimal strategy for Player A to offer an equal 50-50 split.

Return to text.


Figure 2: Dictator Game

Alt Text:
"Game tree diagram illustrating the Dictator Game structure. Player A offers proportion p of the pot to Player B. Player B can accept, resulting in payoffs (1 minus p, p), or reject, resulting in payoffs (1, 0) where Player A keeps everything. The tree shows Player A's initial offer decision followed by Player B's accept-reject decision node."
Long Description:
This sequential game tree shows Player A making an offer of proportion p to Player B. Unlike the ultimatum game, if Player B rejects the offer, Player A still receives the entire pot rather than both players receiving nothing. This structure isolates Player A's other-regarding preferences since Player B's rejection does not punish Player A.

Return to text.


Figure 3: Distribution of Pie Game Outcomes

Alt Text:
"Eight density plots showing the distribution of offers for different language models in the Pie Game. Each subplot shows a model name (Gemma 3, Mistral v0.3, Mistral Small 3.1, Mistral Small 3.2, OLMo 2, Phi 4, Phi 4 Reasoning, Phi 4 Reasoning Plus) with kernel density estimates of offer proportions on the x-axis (0 to 1) and density on the y-axis. All models show strong peaks centered around 0.5, indicating preference for equal splits.”
Long Description:
The figure contains eight panels arranged in a grid, one for each language model tested. All distributions are highly concentrated around p equals 0.5, with very narrow spreads. This indicates that regardless of model architecture or size, LLMs consistently choose equal splits in the Pie Game, which aligns with both self-interest and fairness considerations in this game structure.

Return to text.


Figure 4: Distribution of dictator game outcomes

Alt Text:
"Eight density plots showing the distribution of offers for different language models in the Dictator Game. Models shown are Gemma 3, Mistral v0.3, Mistral Small 3.1, Mistral Small 3.2, OLMo 2, Phi 4, Phi 4 Reasoning, and Phi 4 Reasoning Plus. Unlike Figure 3, these distributions show more variation. Most models display bimodal or concentrated distributions around 0.5, while Gemma 3 shows concentration near 0."
Long Description:
This eight-panel figure reveals diverse behavioral patterns across models in the Dictator Game. Most models show peaks near 0.5 indicating fairness preferences, with some showing bimodal distributions suggesting context-dependent responses. Gemma 3 is notably different, showing a strong concentration near 0, indicating self-interested behavior. Several models display multi-modal patterns suggesting sensitivity to prompt framing.

Return to text.


Figure 5: Implied utilities for various models in the dictator game.

Alt Text:
"Line graph showing implied utility functions for eight language models based on Fehr-Schmidt inequality aversion estimates. The x-axis shows offer proportion p from 0 to 1, and the y-axis shows utility. Multiple colored lines represent different models (Llama 4 Scout, Llama 4 Maverick, Mistral v0.3, Mistral Small 3.1 and 3.2, OLMo 2, and Phi 4 variants). Most models show utility peaks around p equals 0.5, except Llama 4 Maverick which peaks near p equals 0. Gemma 3 results are excluded because Gemma 3 responses to the dictator game were consistently 0.0 and therefore parameter estimates were not obtainable."
Long Description:
The graph displays estimated utility functions based on structural Fehr-Schmidt parameters. Most models exhibit inverted-U shapes peaking near equal splits (p=0.5), demonstrating inequality aversion. Llama 4 Maverick uniquely shows highest utility near p=0, indicating self-interested preferences. The varying slopes and curvatures reflect different degrees of greed aversion (below p=0.5) and envy aversion (above p=0.5) across models. Gemma 3 results are excluded because Gemma 3 responses to the dictator game were consistently 0.0 and therefore parameter estimates were not obtainable

Return to text.


Figure 6: Persona responses to the dictator game for Gemma 3 by college major

Alt Text:
"Density plot showing Gemma 3 model's dictator game offers grouped by application of personas related to college major. The x-axis lists different academic fields (Business, Education, Arts and Humanities, STEM, STEM Related), and the y-axis shows density estimates. Most majors show a peak at approximately 0.25, with Education, Stem, and Business personas displaying a bimodal distribution with another peak centered around the self-interested value of 0.0.” Long Description:

This density plot demonstrates how assigned personas with different educational backgrounds influence Gemma 3's offers. All personas can shift the distribution to a more equitable division of the pot, though the personas related to Education, STEM, and Business are more likely to choose the self-interest optimizing division of the plot than the personas associated with other majors.

Return to text.


Figure 7: Dictator game outcomes by prompt. Responses to prompts vary across models.
OLMo 2
Phi 4 Reasoning

Alt Text:
"Two panels showing density distributions of dictator game offers under different prompt framings. Top panel shows OLMo 2, bottom panel shows Phi 4 Reasoning. Each panel contains multiple overlapping density curves in different colors representing prompt variations: Asset Game, Baseline, Currency Exchange, FOREX, Landlord, and Market Game. Distributions show varying peaks between 0 and 0.6."
Long Description:
These dual-panel density plots illustrate how prompt masking affects offer distributions. Different prompt framings (colors) shift the distributions in systematic ways. For both models, baseline prompts tend to produce higher offers (peaks near 0.5), while FOREX and currency exchange prompts shift distributions toward lower offers (peaks near 0). This demonstrates that recontextualizing the problem reduces fairness concerns and increases self-interested behavior.

Return to text.


Figure 8: Llama 4 (Maverick and Scout) dictator game responses by prompt. Responses to FOREX and currency exchange versions collapse to 0.0.
Llama 4-Scout
Llama4-Maverick

Alt Text: "Two density plot panels comparing Llama 4 Maverick (top) and Llama 4 Scout (bottom) responses across different prompt masks. Multiple colored curves represent different prompts: Asset Game, Baseline, Currency Exchange, FOREX, Landlord, and Market Game. Maverick shows strongest peaks near 0 for most prompts, while Scout shows more variation with some prompts producing peaks near 0.5." Long Description: This figure highlights the different behaviors between two Llama 4 model sizes. Maverick (the larger model) shows strong preference for low offers across most prompts, with pronounced peaks near p=0. Scout (the smaller model) displays greater sensitivity to prompt framing, with bimodal distributions for some prompts, alternating between self-interested (p≈0) and fair (p≈0.5) offers. The FOREX and currency exchange prompts collapse both models toward p=0.

Return to text.


Figure 9: Distribution of responses to the first­person Dictator game at varying levels of control vector strength
Phi 4
Mistral 3.2 Small
Gemma 3
OLMo2

Alt Text:
"Four-panel figure showing kernel density plots for Phi 4, Mistral 3.2 Small, Gemma 3, and OLMo 2. Each panel displays multiple overlapping density curves in different colors representing control vector coefficients from negative 3 to positive 3. For Phi 4 and Mistral, negative coefficients shift distributions toward 0.5 (fairness) while positive coefficients shift toward 0 (self-interest). Gemma 3 and OLMo 2 show more complex patterns."
Long Description:
This figure demonstrates control vector effectiveness across four models. The control vector coefficient acts as a steering parameter: negative values (blue/purple curves) push toward fairness (peaks at 0.5), positive values (yellow/red curves) push toward self-interest (peaks at 0). Phi 4 and Mistral Small 3.2 show clear, predictable responses. Gemma 3 shows initial concentration at 0 that can be shifted toward fairness with extreme negative coefficients. OLMo 2 shows more variable responses across the coefficient range.

Return to text.


Figure 10: Distribution of LLM Responses to the First­Person Dictator Game Across Varying Levels of Control Vector Strength
Phi 4
Mistral 3.2 Small
Gemma 3
OLMo 2

Alt Text:
"Four-panel figure with kernel density plots for Phi 4, Mistral 3.2 Small, Gemma 3, and OLMo 2, showing response distributions to the baseline dictator game under different control vector intensities. Multiple colored curves represent coefficients from negative 2 to positive 2. Patterns show systematic shifts in offer distributions as control vector strength changes, with models responding differently to steering attempts."
Long Description:
Similar to Figure 9 but applied to the baseline dictator game prompt. This figure shows that control vector steering remains effective even with the standard dictator game framing, though the magnitude of shifts varies by model. Most models show increased concentration around p=0 with positive coefficients and shifts toward p=0.5 with negative coefficients, confirming that control vectors can reliably adjust other-regarding preferences in simple allocation tasks.

Return to text.


Figure 11: Fraction of Offers Accepted $$<b$$ vs model size

Alt Text:
"Scatter plot showing the relationship between model size (x-axis, in billions of parameters) and fraction of acceptances (y-axis, from 0 to 0.30). Different colored points represent different prompt framings: Asset Game, Baseline, and Market Game. Larger models generally show lower error rates. Model sizes range from approximately 7B to 27B parameters."
Long Description:
This scatter plot evaluates model rationality by measuring how often models accept offers below the unemployment benefit b, which is always irrational. Larger models (right side) make fewer such errors. The smallest model (Mistral v0.3 at 7B) shows error rates above 0.25. Medium and large models (14B-27B) typically show error rates below 0.10. Prompt framing has some effect, but model size is the dominant factor in reducing irrational behavior.

Return to text.


Figure 12: Policy function examples. Successful experiments on top, unsuccessful on bottom.
Phi 4 successes
Gemma-3 successes
Phi 4 Reasoning Plus failures
Mistral v0.3 failures

Alt Text:
"Four-panel figure showing accept-reject policy functions across 10 experiments for different models. Top row shows Phi 4 and Gemma 3 with clear step functions. Bottom row shows Phi 4 Reasoning Plus and Mistral v0.3 with degraded performance. X-axis shows wage offers centered at reservation wage (0). Y-axis shows binary accept/reject decisions. Top panels show clean threshold behavior; bottom panels show excessive noise or always-accept behavior."
Long Description:
This figure contrasts successful (top) and unsuccessful (bottom) reservation wage identification. The top two panels (Phi 4 and Gemma 3) display clear step functions: offers below the reservation wage are consistently rejected (y=0), offers above are accepted (y=1), with minimal trembling-hand errors. The bottom left (Phi 4 Reasoning Plus) shows the same general pattern but with much higher error rates. The bottom right (Mistral v0.3) fails entirely, accepting nearly all offers regardless of wage level.

Return to text.


Figure 13: Fraction of instances $$\overline {w}$$ exists vs model size

Alt Text:
"Scatter plot showing model success rate in establishing identifiable reservation wages (y-axis, 0 to 1.0) versus model size in billions of parameters (x-axis). Three prompt types shown in different colors: Asset Game, Baseline, and Market Game. Larger models achieve success rates above 0.8, while smaller models drop below 0.4. Gemma 3 at 27B parameters achieves success rates near 1.0 for all prompts."
Long Description:
This plot measures how often models exhibit coherent reservation wage strategies that pass all three rationalizability criteria. Success rates increase strongly with model size. The largest model (Gemma 3, 27B) succeeds in nearly 100% of trials across all prompt types. Mid-sized models (14-24B) show success rates of 60-80%. The smallest model (7B) succeeds in less than 40% of cases. Prompt framing has secondary effects, with baseline prompts generally performing slightly better.

Return to text.


Figure 14: Fraction of instances $$\beta $$ exists vs model size

Alt Text:
"Scatter plot showing the fraction of experiments where discount factor β could be numerically estimated (y-axis, 0 to 1.0) versus model size in billions of parameters (x-axis). Three prompt conditions shown: Asset Game, Baseline, and Market Game. Larger models show higher success rates, with Gemma 3 achieving rates near 0.9. Smallest model shows near-zero success across all prompts."
Long Description:
This figure measures the fraction of experiments where an identified reservation wage can be rationalized with a discount factor β under the McCall model. Success requires both a valid reservation wage and a mathematically consistent β that reproduces that wage given the economic parameters. Large models (24-27B) achieve success in 70-90% of cases. Medium models (13-14B) show more variation (40-80%). The 7B model essentially fails to produce rationalizable behavior, with success rates near zero.

Return to text.


Figure 15: Average $$\beta $$ vs model size

Alt Text:
"Scatter plot showing estimated discount factor β (y-axis, 0 to 1.0) versus model size in billions of parameters (x-axis). Three prompt types indicated by color: Asset Game, Baseline, and Market Game. Most estimates fall between 0.2 and 0.8, with considerable variation. No clear trend with model size. Baseline prompts tend to produce higher β estimates than Asset or Market Game prompts."
Long Description:
This plot displays average estimated patience (β) across models and prompts for cases where β could be recovered. Unlike previous metrics, model size does not show a clear relationship with β magnitude. Instead, prompt framing dominates: baseline prompts (recognizable as McCall problems) produce higher β values (0.5-0.8), suggesting models draw on prior knowledge of the literature. Asset and Market Game prompts produce lower β values (0.2-0.6), potentially reflecting more direct preference revelation uncontaminated by memorized solutions.

Return to text.


Figure 16: StDev $$\beta $$ vs model size

Alt Text:
"Scatter plot showing standard deviation of estimated β values (y-axis, 0 to 0.5) versus model size in billions of parameters (x-axis). Three prompt framings shown in different colors. Most models display standard deviations between 0.2 and 0.3, indicating substantial within-model variability across experiments. No strong relationship between model size and consistency."
Long Description:
This figure reveals that β estimates show high within-model variance regardless of model size. Nearly all models have standard deviations of 0.2-0.3, indicating that estimated patience varies substantially across different economic parameter configurations even within the same model-prompt combination. This suggests that in the sequential McCall task, preferences are less stable and more context-dependent than in the simpler dictator game, where structural parameters were highly consistent.

Return to text.


Figure 17: Successful and unsuccessful application of control vector (Risk) on estimated patience ($$\beta $$)

Alt Text:
"Two scatter plots showing β estimates (y-axis, 0.6 to 0.85; 0.88 to 0.96) versus control vector coefficient (x-axis, negative 1 to positive 1; negative 10 to positive 10). Left panel shows Phi 4 with a clear relationship: higher coefficients produce lower β values from 0.5 to 0.85. Right panel shows Gemma 3 with no relationship: β remains constant near 0.85 regardless of coefficient."
Long Description:
These contrasting panels illustrate variable control vector effectiveness in the McCall search task. Phi 4 (left) responds systematically to the risk-based control vector: positive coefficients reduce patience (β≈0.5), negative coefficients increase patience (β≈0.85), demonstrating successful steering. Gemma 3 (right) is unresponsive: β remains constant around 0.85 across the full range of coefficients, indicating that its preferences in this task are resistant to control vector manipulation despite showing high baseline rationality.

Return to text.


Figure 18: Response to classic dictator at varying intensities of control vector coefficient.

Alt Text:
"Line graph showing mean offer proportion (y-axis, 0 to 0.6) versus control vector coefficient (x-axis, negative 1.5 to positive 1.5) for four models: Phi 4, Mistral 3.2 Small, Gemma 3, and OLMo 2. Lines are smoothed. All models show declining offers as coefficient increases. Phi 4 and Mistral show steepest declines."
Long Description:
This smoothed line plot summarizes the relationship between control vector strength and model offers in the dictator game. All models exhibit the expected pattern: negative coefficients (left) increase fairness preferences producing higher offers, positive coefficients (right) decrease fairness producing lower offers. The rate of change (slope) varies by model, with Phi 4 and Mistral showing the strongest responses. The smoothing reveals generally continuous relationships despite some underlying discontinuities.

Return to text.


Figure 19: Response to FOREX game at varying intensities of control vector coefficient. Responses smoothed. Unsmoothed version in Appendix.

Alt Text:
"Line graph showing mean offer proportion (y-axis, 0 to 0.4) versus control vector coefficient (x-axis, negative 1.5 to positive 1.5) for five models in the FOREX-framed task. Smoothed lines show Phi 4, Phi 4 reasoning, Phi 4 reasoning plus, Gemma 3, and OLMo 2 instruct. All models show lower baseline offers than in Figure 18. Response patterns are less pronounced, with smaller overall shifts across the coefficient range."
Long Description:
This figure parallels Figure 18 but applies to the FOREX-masked dictator game, which naturally induces more self-interested behavior. Baseline offers are lower across all models (y-axis range 0-0.3 versus 0-0.6 in Figure 18). Control vector effects are still present but diminished: the range of offers across coefficients is compressed. This suggests that strong contextual framing (FOREX as trading problem) partially crowds out control vector steering, though directional effects remain consistent.

Return to text.