Whose numbers? What to know before using federal data for Native and rural communities

View of Laguna, New Mexico.

View of Laguna, New Mexico.

Every 10 years the US Census attempts to count everyone in the country, and for the decade that follows, this count quietly steers the allocation of federal dollars. Census-derived data informs the distribution of over $600 billion a year in federal funds for programs like Medicaid, SNAP, Indian Housing Block Grants, and more. If you’re a researcher, grant writer, or program staffer working for or with a Native or rural community, you’re probably using these numbers to help describe a need, justify a grant proposal, or evaluate a program outcome.

Federal datasets like the US Census are not only the easiest to access, but for many communities, they may also be the only available data that speaks to various important issues. Unfortunately, for Native and rural communities, these data are often woefully inaccurate. They undercount true population numbers, define people in ways that communities may not readily recognize, and a recent methodological change implemented for the 2020 Census may have even exacerbated some of these issues. This post explores what’s going wrong and what steps should be taken when working with these datasets considering their problems.

Who gets missed

Despite the Census’ goal of counting everyone, some people and places are reliably missed. Such “hard-to-count” populations tend to combine two features: a rural location and a predominant minority population. This includes Black communities in the rural south, Hispanic communities in the rural Southwest, residents of deep Appalachia, migrant farmworkers, and reservation-based Native communities. In one analysis of the hardest-to-count counties in the US, all 12 counties identified either had mostly Native residents or were rural.

In the 2010 Census, Native Americans living on reservations had a net undercount of 4.9%, the highest of any racial or ethnic group in the nation. Researchers have attributed this to various factors including remote geography, spotty internet access, and a well-earned, historically grounded distrust of government data collection.

Further, undercounting is just part of the problem. The Census records individuals’ self-identified racial and ethnic identity, which is not the same as tribal citizenship, the criteria for which is determined independently by each tribe. Thus, Census data does not provide a reliable accounting of tribal citizenship figures. There’s also the issue of defining “Indian County,” which researchers have pointed out is done in at least nine different ways, including the use of legal land boundaries, Census tribal statistical areas, or simply lumping everyone who self-identifies as “American Indian or Alaska Native” together. The answers to many key questions of interest will vary greatly depending on which definition is used. For example, according to one researcher, only roughly 18.5% of people who identify as “American Indian or Alaska Native” — either alone or in combination with another Census racial category — live in tribal areas. Among those who do live in tribal areas, unemployment measured by the 2010 Census was 10.9%, while among those not living in tribal areas the figure was 6.6% (both of which are likely much lower than the real rates of unemployment). Thus, statements about things like “Native unemployment” or “Native homeownership” can change drastically depending on how the population is defined. Geography adds additional difficulty. Reservations are often “checkerboarded” with federal, state, and private land, and many federal geographic units don’t map neatly onto tribal boundaries.

A Laguna Pueblo tribal member participates in a housing-focused community listening session in late 2024.

A Laguna Pueblo tribal member participates in a housing-focused community listening session in late 2024.

The 2020 privacy change

The US Census Bureau adopted a new method called “differential privacy” for the 2020 Census in response to ever-evolving privacy threats stemming from modern computing. In short, this involves the addition of statistical noise (i.e., small additions or subtractions in the data) to reduce the risk that a third party could reidentify any particular person or household. While the change helps ensure that Census data are protected from re-identification effects, there’s a significant tradeoff in accuracy, especially for small and hard-to-count populations.

Comparing the published 2010 data with 2010 demonstration data put out by the Census Bureau to illustrate the effects of the new methodology, researchers found that while county-level totals and non-Hispanic white counts held up well, counts and growth rates for Black, Hispanic, and Native populations showed significant discrepancies. These discrepancies were also more pronounced in more rural counties, and Native populations fared the worst. The researchers argue that the 2020 Census numbers for such populations are likely affected in the same way, and should thus be treated as noisier and less accurate than they may appear.

A history of extraction

Distrust of data collection in Native communities is rooted in a deep history of extraction and exploitation. There exists, as some researchers put it, a “paradox of scarcity and abundance” as extensive data are collected about Native peoples and communities but less often by or for them. Federal policies of removal, relocation, and assimilation have pushed tribes to depend on outside sources for information about their own communities, despite the fact that Native communities have always relied on the creation and stewardship of their own data to live with the land and pass along traditions and oral histories through the generations. Unfortunately, data collection done by outsiders all too often serves the aims of those collecting data more than the aims of the communities themselves. This pattern has taken new forms in the era of data mining and so-called “big data,” as data extracted from Native communities have been used to market harmful goods and services (e.g., targeting by predatory lenders), inform computing and artificial intelligence systems, and more. This, as one researcher puts it, can be thought of as “digital colonialism.”


Federal policies have pushed tribes to depend on outside sources for information about themselves, despite Native communities relying on the creation and stewardship of their own data for generations. 


For those working within, with, or for Native communities, there are steps that should be taken in light of the ethical and practical concerns that stem from the use of externally collected data. The organizing idea is Indigenous data sovereignty, or the right of Native Nations and communities to govern the collection, ownership, and use of data about their people, lands, and resources. Nations and communities exercise this right through means like data usage codes, review boards, and data-sharing agreements. The CARE Principles (Collective Benefit, Authority to Control, Responsibility, Ethics) crystalize this into guidance that complements the FAIR Principles (Findability, Accessibility, Interpretability, and Reusability) that are common among open-data user circles. While FAIR asks whether data can be used, CARE asks whether it should be used, by whom, and for whose benefit. In practice, this means:

  1. Be clear and say exactly what you are describing. Identify the population for which data is being compiled and reported and stay consistent in the way this population is defined and referred to.

  2. Treat federal numbers as estimates with known blind spots. Uncertainty should be reported, and federal/public data should be complemented by and compared with tribal data (e.g., enrollment records, housing information, community surveys) whenever possible.

  3. Work within each nation or community’s existing processes for approving data collection and analysis, such as review boards, data ownership agreements/requirements, prior review arrangements, and co-authorship agreements.

  4. The limits of any federal or public data used should be clearly described in the write-up.

The bottom line

Federal data is sometimes the only data available to get at questions that are of interest to Native and rural communities. This data can be a useful starting point, but should not be treated as the authoritative truth about such communities. Good practice means maintaining the utmost transparency about its limits and ensuring that the communities behind the numbers are the ultimate decision-makers about how the data is used.



References

  • Carroll, Rodriguez-Lonebear & Martinez (2019). Indigenous Data Governance: Strategies from United States Native Nations. Data Science Journal, 18(31).

  • Junker (2024). Data Mining and Extraction: the gold rush of AI on Indigenous Languages. ComputEL-7.

  • Lopez (2020). Indigenous Data Collection: Addressing Limitations in Native American Samples. Journal of College Student Development, 61(6).

  • Mueller & Santos-Lozada (2022). The 2020 US Census Differential Privacy Method Introduces Disproportionate Discrepancies for Rural and Non-White Populations. Population Research and Policy Review, 41(4).

  • O'Hare (2017). 2020 Census Faces Challenges in Rural America. Carsey School of Public Policy, Brief #131.

  • Payson, S. (2021). Alternative measurements of Indian Country. Monthly Labor Review, 1-18.

  • Roberts, J. S., & Montoya, L. N. (2023, October). In consideration of indigenous data sovereignty: data mining as a colonial practice. In Proceedings of the Future Technologies Conference (pp. 180-196). Cham: Springer Nature Switzerland.

  • Rodriguez-Lonebear, D. (2016). Building a data revolution in Indian country. Indigenous data sovereignty: Toward an agenda, 14, 253-72.

  • Ruggles, S. (2024). When privacy protection goes wrong: how and why the 2020 census confidentiality program failed. Journal of Economic Perspectives, 38(2), 201-226.

 

Share


Related content


Max Van Oostenburg

Max is an associate research director at Sweet Grass who studies how people connect with themselves, their communities, and the world around them. He received his PhD in anthropology from the University of Florida.

Next
Next

The goal of participatory research is a true shift in power