Datasets are available upon request.
However, please review the details on this page before using the data in your research.
[A] Commuting Zones
[B] Metropolitan Statistical Areas
[C] Industry Code
[D] Industry Employment
[E] Industry Trade Exposure
[F] Government Fiscal Data
[G] Population Data
Paper: "Dealing with Fiscal Stress: Cities versus Suburbs"
Why do you include only 709 CZs in your paper instead of 722 CZs, as in the papers by Autor et al. and Feler and Senses?
Ans: Commuting zones (CZs), which represent local labor markets, were developed by Tolbert and Sizer (1996). They use county-level commuting data from the 1990 Census to construct 741 clusters of counties. There are 722 CZs in the contiguous United States.
I exclude the following CZs from my paper because of duplication or other data issues:
1. drop #10600 keep #10700: they have the same largest cities in their CZs, Birmingham city, AL. #10700 is larger than #10600 in terms of population size.
2. drop #11304: District of Columbia has a very different character than other CZs
3. drop #20402: no fiscal data in suburbs in Nantucket County, MA
4. drop #27704: no fiscal data in central city in year 1992 and before
5. drop #30605: no fiscal data in suburbs in Culberson County, TX
6. drop #31304: no fiscal data in suburbs in Mason County, TX
7. drop #31503: no fiscal data in suburbs before 1992
8. drop #32306: no fiscal data in suburbs in Maverick County, TX
9. drop #32603: no fiscal data in suburbs in Baylor County, TX
10. drop #34101 - #34115: CZs are in Alaska state
11. drop #34701 ,34702 ,34703 ,35600: CZs are in Hawaii state
12. drop #34306: no fiscal data in suburbs in 30033 (fips_state_county) Garfield County, MT
13. drop #34307: no fiscal data in central city in year 1997 and before
14. drop #37902: no fiscal data in suburbs in 32021 (fips_state_county) Mineral County, NV
15. drop #39301: no fiscal data in suburbs in 53055 (fips_state_county) San Juan County, WA
Papers: "Rent Capture by Central Cities," "Metropolitan Fragmentation and Transportation Investment"
1. Which geographic unit is larger: commuting zones or Metropolitan Statistical Areas?
Ans: Metropolitan Statistical Areas (MSAs) are defined by the Office of Management and Budget (OMB) and differ from commuting zones in how their boundaries are constructed. MSAs include at least one county and only consider places with populations above 50,000.
There is no definitive answer because CZs are larger than MSAs in some cases, while MSAs are larger than CZs in others.
2. How do you define suburbs in the first ring (mc = 2), second ring (mc = 3), or third ring (mc = 4)? Why do you separate suburbs into different rings? How does this approach differ from using an aggregated measure?
Ans: I mainly follow the 1990 Census definitions when constructing the MSA data.
) Download the county shapefile: https://www.census.gov/geographies/mapping-files/time-series/geo/carto-boundary-file.1990.html#list-tab-1556094155
) Import the shapefile into ArcGIS or QGIS and duplicate the layer if you want to create a map with two layers.
) Select “Layer” > “Open Attribute Table” to view the data.
) Select “Layer” > “Filter”, then select the county you want to keep.
) Select “View” > “Identify Features.”
) Define counties in the first ring (mc = 2) as those that share a boundary with the main county (mc = 1), where the central city is located.
) Define counties in the second ring (mc = 3) as those that share a boundary with a first-ring county (mc = 2) but do not share a boundary with the main county (mc = 1).
The purpose of separating suburbs into different rings is to examine suburban local governments in different economic environments. Some suburbs have relatively small populations, while others are large enough to function as twin cities or satellite cities.
Aggregating all suburban local governments into a single unit may average out local characteristics. A possible advantage of aggregation is that it internalizes coordination and competition among suburbs, while a possible disadvantage is that it may introduce aggregation bias.
I mainly focus on MSAs in the top population quartile (approximately 70 MSAs) based on the 1990 Census definition. Most of these MSAs have first-ring suburbs, while only a small number (approximately 10 MSAs) have both first- and second-ring suburbs. Only MSAs in Virginia have first-, second-, third-, and fourth-ring suburbs.
Papers: "Dealing with Fiscal Stress: Cities versus Suburbs", "Rent Capture by Central Cities"
1. Which industry classification system do you primarily use, and why?
Ans: I primarily use four-digit SIC industries. Pierce and Schott (2009) assign 10-digit HS products to four-digit SIC industries, ensuring that each of the 397 manufacturing industries is matched to at least one trade code.
2. What challenges do you face when merging data from different coding systems?
Ans: County Business Patterns (CBP) reports employment data by county and industry using six-digit NAICS codes in 2000. However, for 1980 and 1990, CBP reports industry-level employment data using four-digit SIC codes. Using the Census “bridge” file, we can construct a weighted crosswalk between the two classification systems. (Details are available upon request.)
Papers: "Dealing with Fiscal Stress: Cities versus Suburbs"
1. Why are the Bartik instruments you construct based on the 1980 employment share instead of the 1986 or 1989 employment share as a proxy for the 1990 local employment share?
Ans: County Business Patterns (CBP) provides the most disaggregated employment data for Census years, with more detailed industry-level information than is available for other years.
Papers: "Dealing with Fiscal Stress: Cities versus Suburbs"
1. How are trade shocks defined? What is the difference between Autor et al. (2013) and Autor et al. (2021)?
Ans: Generally, trade shocks are measured as changes in import value per worker, import penetration, or import exposure. However, Autor et al. (2021) define trade shocks as changes in import value divided by domestic industry absorption (U.S. industry shipments plus net imports).
2. Why does the analysis of the evolving impact of trade shocks on local variables use different time horizons for the X and Y variables? More specifically, why is the specification written as follows?
Yi,t+h=a1+b1×Shock2000−2012+Xit′b2+ei,t+h,h=1,…,12
Ans: The trade shock is measured over the 2000–2012 period, beginning one year before China joined the WTO and extending beyond the culmination of the trade shock around 2010. Workers may take time to migrate or find employment in the non-manufacturing sector, so the full adjustment along these margins may only become apparent after the trade shock reaches its peak.
3. What about spillover effects across industries? Chinese import shocks may not always have adverse effects on local areas.
Ans: Acemoglu et al. (2016) use U.S. input-output data to construct supplier (upstream) and customer (downstream) import-exposure shocks for both manufacturing and non-manufacturing industries. When customer industries are directly affected by trade shocks, related industries experience adverse employment effects. However, there is no evidence of significant employment changes in industries whose suppliers are directly affected by trade shocks.
Papers: ALL
1. Where do you obtain government fiscal data? Are alternative sources available?
Ans: I obtain government fiscal data from the U.S. Census Bureau’s Annual Survey of State and Local Government Finances.
An alternative source is:
Pierson, K., Hand, M., and Thompson, F. (2015). The Government Finance Database: A Common Resource for Quantitative Research in Public Financial Analysis. PLoS ONE. doi: 10.1371/journal.pone.0130119.
Most local governments report their fiscal data every five years, so I fill in missing values using one of two methods: linear interpolation or a constant growth rate.
Papers: ALL
1. Besides cities in the Rust Belt, which U.S. cities are experiencing population decline?
Ans: My recent research finds that many declining cities began losing population around 1950.
Chicago (msa_sc = 15): The population peaked in 1950 and declined until 1990. Since 1990, it has fluctuated, and the population in 2020 was close to its 1990 level.
Baltimore (msa_sc = 19): The population peaked in 1950 and has declined since then.
Detroit (msa_sc = 21): The population peaked in 1950 and has declined since then. By 2020, its population had fallen to approximately one-third of its 1950 level.
St. Louis (msa_sc = 24): The population peaked in 1950 and has declined since then.
Cincinnati (msa_sc = 30): The population peaked in 1950 and has declined since then.
Cleveland (msa_sc = 31): The population peaked in 1950 and has declined since then.
Toledo (msa_sc = 33): The population peaked in 1970 and has declined since then.
Philadelphia (msa_sc = 37): The population peaked in 1950 and has declined since then.
Pittsburgh (msa_sc = 38): The population peaked in 1950 and has declined since then.
Milwaukee (msa_sc = 50): The population peaked in 1960 and has declined since then.
Buffalo (msa_sc = 61): The population peaked in 1950 and has declined since then.
Birmingham (msa_sc = 64): The population peaked in 1960 and has declined since then.
Rochester (msa_sc = 66): The population peaked in 1950 and has declined since then.
Akron (msa_sc = 91): The population peaked in 1960 and has declined since then.
Dayton (msa_sc = 93): The population peaked in 1960 and has declined since then.
Shreveport (msa_sc = 103): The population peaked in 1980 and has declined since then.
Mobile (msa_sc = 105): The population peaked in 1960 and has declined since then.
Jackson (msa_sc = 106): The population peaked in 1980 and has declined since then.