Abstract:
Keywords: housing booms, beliefs, transaction data.
JEL Classification: D14, D91, R21, R31
Optimistic beliefs and looser credit conditions are two highly studied drivers of the U.S. housing boom of the 2000s.1 While these drivers are often studied in isolation, which allows for clean interpretation of mechanisms, studying them as complementary can allow for propagation dynamics arising from their interaction.2 However, studying their interaction is challenging due to limited data on beliefs prior to 2007.3 This paper finds that beliefs were higher in relatively lower-priced, rather than higher-priced, houses in the 2000s for all locations studied. Since buyers of lower-priced houses are more sensitive to credit conditions Fidelman and Tapak, 2026, this paper’s evidence of higher expected capital gains in the lower-priced housing segment points to an interaction of beliefs and credit conditions.4 Interpreted more broadly, this evidence suggests that beliefs and credit conditions were complementary, rather than competing, drivers of the U.S. housing boom of the 2000s.
To investigate how beliefs vary across housing price segments, this paper estimates a statistical model of price changes developed by Landvoigt et al., 2015 using Zillow, 2026 ZTRAX transaction data on repeat housing sales. Transaction-level data provides repeat sales prices of the same property, which, in turn, allows for the estimation of a common component of expected capital gains and a cross-sectional dispersion component across types of houses segmented by price. This model assumes a one-dimensional quality index so that the sale price of a house fully reflects its quality. Consequently, as explained by Landvoigt et al., 2015, any statistical model of price changes can give rise to an expected price change and thus expected capital gains, which can, in turn, be used to proxy for beliefs.
By expanding the analysis of Landvoigt et al., 2015 beyond San Diego, CA to include Phoenix, AZ and Cleveland, OH, this paper adds new evidence to the dearth of data on beliefs in the 2000s. Because the U.S. housing boom of the 2000s varied in timing and magnitude across metropolitan statistical areas (MSAs), as documented by Ferreira and Gyourko, 2023, estimating beliefs for multiple MSAs is important for a more comprehensive account of the episode. For this reason, I study three MSAs that capture characteristics of different submarkets, but all show higher expected capital gains in relatively lower-priced houses. San Diego is typically characterized by authors like Saiz, 2010 as having a low elasticity of housing supply such that it is difficult—because of geography—to build more houses in response to higher house prices. On the other hand, Phoenix and Cleveland are both characterized as having geography that lends to a higher elasticity of housing supply. However, their fundamentals differ; Phoenix in the 2000s faced rapid housing and population growth and Cleveland depopulation.5
Complementing these within-MSA findings, this paper also estimates substantial variation in expected capital gains over time and across the three MSAs studied, which corroborates the empirical evidence of Soo, 2018.
Transactions of repeat sales of single-family homes for the MSAs of Cleveland, OH; Phoenix, AZ; and San Diego, CA are obtained via Zillow, 2026 ZTRAX assessment and transaction data accessed through the Bureau of Economic Analysis in 2019. Applying the cleaning steps detailed in Appendix A to arm’s length, non-foreclosed sales of residential properties made by owner-occupiers results in over 280,000 transactions of houses that are sold at least once from 1998 to 2008.6 Limiting the sample to repeat sales is essential because the statistical model estimates expected capital gains from the observed sales of the same property.
Expected capital gains of house \(i\) at time \(t\) in MSA \(j\) vary by their current log price \(\log p_t^{i,j}\) where the idiosyncratic shocks \(e^{i,j}_{t+1}\) have mean zero and are such that the law of large numbers holds in the cross section of houses. Because this paper assumes that there is a one-dimensional quality index, quality is reflected one-for-one in the house price at any given time \(t\) and there are no additional controls.7
The statistical model can be written as:
\[\begin{equation} \log p^{i,j}_{t+1}-\log p^{i,j}_{t}=a_{t+1,t}^j+b^j_{t+1,t}\log p^{i,j}_{t}+e^{i,j}_{t+1,t}\tag{1} \end{equation}\]
Estimating the equation above would restrict the sample to houses that sold in successive years, such as a sale in 2000 and 2001, for example. To incorporate information from longer-dated sales, such as a house that sold in 2000 and again in 2003, for example, the above equation is iterated forward to capture expected capital gains between years \(t+k\) and \(t\) for each \(t=1998,...,2007\) and \(k \in [1,2008-t]\). \[\begin{equation*} \log p^{i,j}_{t+k}-\log p^{i,j}_{t}=a^j_{t+k,t}+b^j_{t+k,t}\log p^{i,j}_{t}+\epsilon^{i,j}_{t+k,t} \end{equation*}\] For \(k\geq2\), the above equation can be written in terms of \(a_{t+1,t}\) and \(b_{t+1,t}\) . \[\begin{equation} log p^{i,j}_{t+k}=a^{i,j}_{t+k,t+k-1} + \sum_{m=t+2}^{t+k}\prod_{\ell=m}^{t+k} (1+b^{j}_{\ell,\ell-1})a^{j}_{m-1,m-2} +\prod_{\ell=t+1}^{t+k}(1+b_{\ell,\ell-1})\log p_t^{i,j}\tag{2}\end{equation}\]
Equations (1) and (2) are estimated separately for each MSA \(j\in\{\text{Cleveland}, \\ \: \text{Phoenix}, \: \text{San Diego}\}\). Estimation is for \(t=1998,...,2007\) via two-step GMM with a robust weight matrix. The first-step residuals are used to estimate the weighting matrix for the second step, which ensures that the estimator is robust to heteroskedasticity. Estimates \(\hat{a}^j_{t+k,t}\) and \(\hat{b}^j_{t+k,t}\) can be obtained by adding/multiplying \(\hat{a}^j_{t+1,t}\) and \(\hat{b}^j_{t+1,t}\), respectively.
The slopes \(b^j_{t+1,t}\) are the coefficients of interest and test for cross-sectional dispersion in expected capital gains across house-price segments for each MSA \(j\) for any given pair of years \(t+1\) and \(t\), conditional on the full sample of transactions between years \(t+k\) and \(t\). If \(\hat{b}^j_{t+1,t}=0\) then there is no cross-sectional dispersion and all houses in that MSA have the same expected capital gains regardless of their price. If \(\hat{b}^j_{t+1,t}>0\), then expected capital gains are relatively higher for more expensive houses in MSA \(j\), which are those houses that have a higher sales price at year \(t\). If \(\hat{b}^j_{t+1,t}<0\), relatively cheaper houses in MSA \(j\) have higher expected capital gains. Because Fidelman and Tapak, 2026 find that households in lower-priced housing segments tend to be more sensitive to credit conditions, testing if \(\hat{b}^{j}_{t+1,t}=0\) can provide insights on the interaction of beliefs and credit conditions. The intercepts \(a^j_{t+1,t}\) are the average expected capital gain. Figure 1 shows the GMM estimates of the coefficients from equations (1) and (2) as the markers and 95% confidence intervals as the shaded bands for Cleveland, Phoenix and San Diego. Estimates shown are those from one year \(t\) to the next year \(t+1\) such that \(k=1\). While these can be added/multiplied to obtain those that correspond to \(k\geq2\), those for \(k=1\) can be more readily compared to actual 12-month percentage change in house prices as shown in figure 2.8
The top panel 1a shows that expected capital gains were relatively higher for lower-priced houses for all three MSAs shown until about 2005, as shown by the estimates of the cross-sectional dispersion of expected capital gains across house price segments \(\hat{b}^j_{t+1,t}\). These estimates show that a 1 percent less expensive house (relative to the average house) was expected to have at most a 0.1 percentage point higher expected capital gain.
Comparing across MSAs shown, panel 1a points to lower-priced houses having relatively higher expected capital gains in San Diego than in Phoenix or Cleveland, on average. Because the housing supply elasticity is lower in San Diego than in Phoenix or Cleveland (according to Saiz, 2010, Saiz, 2010), increased demand for less expensive houses was more likely to result in relatively higher price increases, and thus capital gains.9 Conversely, Phoenix’s near-zero dispersion suggests relatively uniform expected capital gains across house-price segments. This can be attributed to, in part, a high share of speculative investors who—per Gao et al., 2020—tended to be insensitive to credit conditions.10
Figure 2: 12-month percentage change in house prices for select
metropolitan statistical areas, percentage points. Dark shaded bands are
NBER recessions and the light shaded banded between the two vertical
lines is the U.S. housing boom period defined as 1998 to 2006 in this
paper. Source: S&P Cotality Case-Shiller Home Price Indices,
National Bureau of Economic Research.
By construction, the estimated average expected capital gains, \(\hat{a}^j_{t+1,t}\) shown in the bottom panel 1b, closely track the 12-month percentage change in house prices shown in figure 2. This alignment helps validate that the estimated data generating process in equations (1) and (2) correctly reflects key attributes about regional house prices—specifically the early peak for San Diego, the later arriving rapid surge for Phoenix, and virtually no boom for Cleveland. Notably, expected capital gains in Phoenix and Cleveland were similar and remained below those of San Diego until 2003, which likely reflects lower buyer incomes in these regions, as noted by Ferreira and Gyourko, 2023. However, after 2003, expected capital gains in Phoenix and Cleveland diverge sharply, which is likely due to the influx of housing investors in the former, but not the latter.
It is also worth noting that figure 1 shows substantial time variation in estimates of both average and cross-sectional expected capital gains for each MSA. By the start of the bust in 2007, the positive values of \(\hat{b}^j_{t+1,t}\) for all MSAs shown indicate that relatively more expensive houses were holding their value than less expensive houses, while the negative value of \(\hat{a}^j_{t+1,t}\) indicates that average expected capital gains were negative. This time variation is important for understanding the evolution of beliefs throughout various stages of a boom-bust cycle.
The boom in national U.S. house prices is largely attributed to optimistic beliefs and looser credit conditions. Because data on beliefs prior to 2007 is limited, studying the interaction of beliefs and credit conditions is challenging. Estimating the statistical model of Landvoigt et al., 2015 on transaction-level housing data for the MSAs of Cleveland, OH; Phoenix, AZ; and San Diego, CA shows that lower-priced houses had higher expected capital gains than their higher-priced counterparts. Because buyers of lower-priced houses tend to be more sensitive to looser credit conditions Fidelman and Tapak, 2026, and these houses had higher expected capital gains, this paper documents a pattern that supports an interaction of credit conditions and optimistic beliefs about future housing.
This section gives detailed cleaning steps of the Zillow, 2026 ZTRAX transaction and assessment data to obtain a sample similar to that of Landvoigt et al., 2015 for the MSAs of Cleveland, OH; Phoenix, AZ; and San Diego, CA. The Zillow, 2026 ZTRAX data is a panel of housing transactions and the cleaning steps can be grouped by those related to deeds, characteristics, and outliers.
First, to obtain housing market transactions, deeds
(documenttype) that are not arm’s length transfers of homes
are dropped as are other non-arm’s length deeds such as those indicating
a partial sale of a house. Following Landvoigt et al., 2015, I keep only grant deeds
(GRDE), condo deeds (CDDE), corporate deeds
(CPDE), and individual deeds (IDDE). Because
the list of deeds denoting arm’s length transactions is more exhaustive
for Cleveland, OH and Phoenix, AZ, I delete entries for
documenttype that are conservator’s deed
(CVDE), deed in lieu of foreclosure (DELU),
gift deed (GFDE), intrafamily transfer (INTR),
partnership deed (PTDE), personal representative’s deed
(PRDE), sheriff’s deed (SHDE), trustee’s deed
(TRFC). Deeds that remain have values for
documenttype that include administrator’s deed
(ADDE), agreement of sale (AGSL), bargain and
sale deed (BSDE), condominium deed (CPDE),
court order/action (COCA), corporation deed
(CPDE), correction deed (CRDE), deed
(DEED), fiduciary deed (FDDE), guardian’s deed
(GDDE), grant deed (GRDE), individual deed
(IDDE), joint tenancy deed (JTDE), land
contract (LDCT), other (OTHR), quitclaim deed
(QCDE), re-recorded deed (RRDE), tax deed
(TXDE), and warranty deed (WRDE).
Second, deeds are dropped based on buyer or house characteristics.
Deeds without a latitude or longitude are dropped as are deeds that
transfer multiple parcels as identified by the APN number. Second homes
and trailers are dropped as are foreclosed properties. Buyers that are
not a couple or single person are dropped eliminating buyers that are a
corporation or partnership (CO,PT), a trust
(FT,IT,LV,RL,RT,TE),
or a beneficiary (BF).11
Single-family homes are denoted by the propertylanduse
variable from the assessment data12 and observations that
are kept include single-family residences (RR101),
condominiums (RR106)13, cooperatives
(RR107), row houses (RR108), planned unit
developments (RR109), bungalows (RR113), zero
lot lines (RR114), manufactured, modular and prefabricated
homes (RR115), patio homes (RR116), garden
homes (RR119), landominiums (RR120), and
inferred single-family homes (RR999). Dropped observations
have a propertylanduse variable from the assessment data
equal to rural residences including farms/productive land
(RR102), mobile homes (RR103), residential
common areas (RR110), time shares (RR111),
seasonal, cabin, vacation residences (RR112), residential
parking garages (RR117), and other improvements
(RR118). Observations from the transaction data are also
dropped such as those where the propertyusestndcode
variable equals agricultural (AG), apartment
(AP), commercial (CM), mobile homes
(MB), mixed use (MX), unimproved
(UL), multifamily (MF).
Lastly, to control for outliers, transactions with prices below $15,000, combined loan-to-value ratios (first plus second mortgage) above 120 percent, and annualized capital gains above 50% are dropped. To avoid the influence of house flipping, all pairs of sales that are less than 180 days apart are dropped. Some properties that were sold twice in the same year, but more than 180 days apart remain in the sample. If the property was sold more than once in the same year, then these transactions are dropped. If the property was sold in another year then the earlier sale date in the year is dropped. Given that house prices are rising throughout this period, keeping the later sale should bias capital gains downward. Given the high house price growth observed at the MSA-level in Phoenix, AZ, robustness checks were run to increase the annualized capital gain threshold from 50% to 60% and 70% with little change to the estimates. Similarly, second family homes were included in a separate robustness check with little alteration to the estimates.
Table 1 compares the estimates of average and cross-sectional capital gains for San Diego, CA to those of Landvoigt et al., 2015. Overall, the estimated coefficients from equations (1) and (2) resemble those of Landvoigt et al., 2015. A few minor discrepancies likely arise from the differences in source data used (Trulia vs. Zillow). My sample is larger than theirs with 84,076 repeat sales compared to their 70,315.
Table 1: Upper panel contains estimates from Landvoigt et al., 2015 (LPS) of 70,315 repeat sales in San Diego County during the years 1999-2008 using transaction data from Trulia. The lower panel contains a replication of their estimates using ZTRAX assessment and transaction data for 84,076 repeat sales in San Diego County. The numbers in parentheses are standard errors.
| 2000 | 2001 | 2002 | 2003 | 2004 | 2005 | 2006 | 2007 | |
|---|---|---|---|---|---|---|---|---|
| \(a_{t+1,t}^{LPS}\) | 1.29 (0.04) | 1.41 (0.04) | 1.30 (0.04) | 0.87 (0.05) | 0.60 (0.06) | -0.56 (0.07) | -1.09 (0.10) | -3.18 (0.12) |
| \(b_{t+1,t}^{LPS}\) | -0.093 (0.003) | -0.10 (0.003) | -0.09 (0.003) | -0.05 (0.004) | -0.04 (0.004) | 0.04 (0.01) | 0.07 (0.01) | 0.22 (0.01) |
| \(a_{t+1,t}^{author}\) | 1.04 (0.04) | 1.37 (0.04) | 1.23 (0.04) | 0.71 (0.05) | 0.55 (0.05) | -0.33 (0.07) | -1.84 (0.10) | -3.16 (0.12) |
| \(b_{t+1,t}^{author}\) | -0.07 (0.003) | -0.10 (0.003) | -0.08 (0.003) | -0.04 (0.004) | -0.03 (0.004) | 0.02 (0.005) | 0.13 (0.008) | 0.21 (0.009) |
Aastveit, K. A., Albuquerque, B., and Anundsen, A. K. (2023).
Changing supply elasticities and regional housing booms.
Journal of Money, Credit and Banking,
55(7):1749–1783.
Albanesi, S., DeGiorgi, G., and Nosal, J. (2022).
Credit growth and the financial crisis: A new narrative.
Journal of Monetary Economics, 132:118–139.
Bayer, P., Mangum, K., and Roberts, J. W. (2021).
Speculative fever: Investor contagion in the housing bubble.
American Economic Review, 111(2):609–51.
Case, K. E. and Shiller, R. J. (1987).
Prices of single-family homes since 1970: new indexes for four
cities.
New England Economic Review, September 1987, pages 45-56.
Chinco, A. and Mayer, C. (2016).
Misinformed speculators and mispricing in the housing market.
Review of Financial Studies, 29(2):486–522.
Cox, J. and Ludvigson, S. C. (2021).
Drivers of the great housing boom-bust: Credit conditions, beliefs, or
both?
Real Estate Economics, 49(3):843–875.
Dong, D., Liu, Z., Wang, P., and Zha, T. (2022).
A theory of housing demand shocks.
Journal of Economic Theory, 203:105484.
Duca, J. V., Muellbauer, J., and Murphy, A. (2021).
What drives house price cycles? international experience and policy
issues.
Journal of Economic Literature, 59(3):773–864.
Ferreira, F. and Gyourko, J. (2023).
Anatomy of the Beginning of the Housing Boom across U.S.
Metropolitan Areas.
The Review of Economics and Statistics,
105(6):1442–1447.
Fidelman, T. and Tapak, T. (2026).
Housing markets and macroeconomic shocks.
Working Paper.
Gao, Z., Sockin, M., and Xiong, W. (2020).
Economic consequences of housing speculation.
The Review of Financial Studies,
33(11):5248–5287.
Glaeser, E. L. and Gyourko, J. (2005).
Urban decline and durable housing.
Journal of Political Economy, 113(2):345–375.
Glaeser, E. L., Gyourko, J., and Saiz, A. (2008).
Housing supply and housing bubbles.
Journal of Urban Economics, pages 198–217.
Glaeser, E. L. and Tobia, K. (2007).
The rise of the sunbelt.
Taubman Center Policy Brief.
Graham, J. (2024).
House prices, investors, and credit in the great housing bust.
Working Paper.
Jacobson, M. M. (2024).
Beliefs, aggregate risk, and the u.s. housing boom.
Manuscript.
Johnson, S. (2019).
Mortgage leverage and house prices.
Working Paper.
Kaplan, G., Mitman, K., and Violante, G. (2020).
The Housing Boom and Bust: Model Meets Evidence.
Journal of Political Economy, 128(9).
Kuchler, T., Piazzesi, M., and Stroebel, J. (2023).
Chapter 6 - housing market expectations.
pages 163–191.
Landvoigt, T., Piazzesi, M., and Schneider, M. (2015).
The housing market(s) of san diego.
American Economic Review, 105(4):1371–1407.
Louie, S., Mondragon, J. A., and Wieland, J. (2025).
Supply constraints do not explain house price and quantity growth across
us cities.
Working paper.
Mian, A. and Sufi, A. (2022).
Credit supply and housing speculation.
The Review of Financial Studies,
35(2):680–719.
Oh, H., Yang, C., and Yoon, C. (2025).
Land development and frictions to housing supply over the business
cycle.
Working Paper.
Piazzesi, M. and Schneider, M. (2016).
Housing and macroeconomics.
In Taylor, J. B. and Uhlig, H., editors, Handbook of
Macroeconomics, volume 2, chapter 19. North Holland.
Saiz, A. (2010).
The geographic determinants of housing supply.
The Quarterly Journal of Economics,
125(3):1253–1296.
Soo, C. (2018).
Quantifying sentiment with news media across local housing
markets.
The Review of Financial Studies,
31(10):3689–3719.
Zillow (2026).
Zillow transaction and assessment database (ztrax), united states,
1940-2020.
Inter-university Consortium for Political and Social Research
[distributor], 2026-01-27. https://doi.org/10.3886/ICPSR39652.v2.
buyercode equal to domestic partners (DP),
formerly known as (FK), her husband (HH),
husband and wife (HW), individual (ID),
married man (MM), minor (MN), married person
(MP), married woman (MW), single man
(SM), single person (SP), single woman
(SW), unmarried man (UM), unmarried woman
(UW), widowed (WW) and dropping those with
buyercode equal to affiant (AF), borrower or
trustor in default (BR), estate (ES), executor
(EX), government (borough, city, village, etc.)
(GV), surviving joint tenant (SJ), personal
representative (PR), agent (AG), not provided
(NP). Return to Textpropertylanduse variable from the
assessment data and propertyusestndcode variable from the
transaction data are mostly but not always consistent. I use the
assessment data as the main variable to denote property use because it
has fewer blank observations and is more stable across time. Return to Text