The Public File Is Not the Real File: Understanding What Federal Agencies Withhold When They Release Data
There is a common assumption among researchers new to federal data that the publicly available version of a government dataset represents the full scope of what was collected. This assumption is incorrect in almost every case that matters.
Federal statistical agencies operate under competing obligations. On one side sits the mandate for transparency—the principle that data collected with public funding should be accessible to researchers, journalists, and citizens. On the other side sits a set of legal, administrative, and political constraints: Title 13 confidentiality protections for Census Bureau data, HIPAA restrictions on health records, disclosure avoidance requirements that prohibit releasing any table cell that could be traced back to an individual respondent.
The resolution of that tension, in practice, is a public-use file that has been substantially altered from the internal analytical file the agency itself uses. Understanding the nature and extent of those alterations is not a peripheral methodological concern. It is a precondition for knowing whether your analysis is answering the question you think it is.
Where the Data Goes When It Leaves the Agency
The journey from administrative record or survey response to public dataset involves a series of transformations, each of which trades analytical precision for disclosure protection.
Geographic suppression is among the most consequential. The American Community Survey's public-use microdata—the PUMS files—are among the richest individual-level datasets the federal government produces. But PUMS geography is limited to Public Use Microdata Areas of at least 100,000 residents. Counties smaller than that threshold, and all sub-county geographies, are entirely absent from the public file. For researchers studying rural communities, small metropolitan areas, or neighborhood-level variation, the public PUMS is functionally unusable.
Variable recoding and top-coding affect income, earnings, and age variables across virtually every federal microdata product. The Current Population Survey public-use file top-codes individual earnings at a threshold that has not kept pace with wage growth at the upper end of the distribution, meaning that high-income households are systematically misrepresented in any analysis that relies on the public file for income estimates.
Sample thinning is less visible but equally damaging. Some agencies release public microdata that represents a subsample of the full survey—a 1-in-100 sample where the internal file contains every respondent. This is not always clearly disclosed in the documentation accompanying the public release, and researchers who calculate standard errors from the public file without accounting for the subsample design will produce confidence intervals that are far too narrow.
Synthetic data injection is an increasingly common disclosure avoidance technique. Rather than simply suppressing sensitive values, some agencies replace them with statistically plausible synthetic values generated from the distribution of the observed data. The 2020 decennial Census introduced differential privacy mechanisms that inject calibrated noise into small-area tabulations. The result is a public dataset that is statistically consistent at the national level but may be substantially inaccurate for small geographies or small demographic subgroups.
The Datasets Most Affected
Not every federal dataset is equally compromised in its public release. Researchers should apply heightened scrutiny to the following programs:
The Medical Expenditure Panel Survey public files omit geographic identifiers below the Census region level, making any state-level or metropolitan-area analysis impossible without restricted access. The restricted-use files contain full geographic detail and are available to qualified researchers through the Agency for Healthcare Research and Quality's data enclave.
The Survey of Income and Program Participation public microdata has been subject to multiple waves of variable suppression over successive redesigns. Variables related to asset values, program participation spells, and employer-provided benefit details that appeared in earlier panels have been progressively restricted, making cross-panel longitudinal comparisons structurally unreliable.
The Longitudinal Employer-Household Dynamics program produces some of the most analytically valuable labor market data the federal statistical system generates. The public Quarterly Workforce Indicators are aggregated to the county level with suppression rules that eliminate a substantial share of cells for rural counties and small industries. The underlying linked employer-employee microdata exists only in restricted form.
The National Health Interview Survey public-use file omits state identifiers for states with small populations and suppresses detailed diagnostic codes that appear in restricted files. For health services researchers conducting state-level or condition-specific analyses, the public file frequently cannot support the research design.
How to Access What the Public File Withholds
The pathway to restricted-use data varies by agency but follows a common structure. Researchers typically must demonstrate institutional affiliation, describe the research project and its public benefit, agree to data security requirements, and in some cases conduct analysis within a secure computing environment maintained by the agency rather than on their own systems.
The Census Bureau's Federal Statistical Research Data Centers network is the primary access point for restricted microdata from the ACS, CPS, SIPP, and several other programs. Applications require IRB approval, a project proposal, and a disclosure avoidance review of any outputs before they can be removed from the secure environment. The process is time-consuming—approval timelines of six months to a year are not unusual—but the analytical payoff for research requiring geographic detail or sensitive variables is substantial.
For health data, AHRQ, the National Center for Health Statistics, and the Centers for Medicare and Medicaid Services each maintain separate restricted-access programs with distinct application procedures. Researchers working across health domains should not assume that approval for one program transfers to another.
Interpreting Absence as Evidence
When a variable is suppressed, a geographic unit is collapsed, or a cell is replaced with a synthetic value, the absence is itself analytically meaningful. Small-area suppression in federal datasets is not random—it is concentrated in geographies where populations are small, often rural, often economically marginal, and often demographically distinct from national averages. Research that relies exclusively on public files is therefore systematically biased toward the kinds of places and populations that are large enough to survive disclosure avoidance rules.
For data professionals advising policymakers or producing research intended to inform resource allocation, this is not a technical footnote. It is a structural limitation that shapes which communities are visible in the evidence base and which are not. Knowing where the public file ends—and having a plan for what to do when it does—is as fundamental a research skill as knowing how to run a regression.