announcements

Specifically for meta-information about the site

Map of San Antonio Showing Reliable Estimates by census tract

Comparing 5-year ACS Estimates Over Time

My article on comparing American Community Survey (ACS) estimates over time was published earlier this month. It will take another 6 to 12 months before it’s slotted into a actual volume and issue, but until then it’s available online via the link in this citation:

Donnelly, F. P. (2026). The Feasibility of Comparing 5-Year American Community Survey Estimates Over Time. Journal of the American Planning Association, 1–14. https://doi.org/10.1080/01944363.2026.2709860

Unfortunately it’s not open access; if you don’t have a subscription, you can read the article in the publisher’s e-pub platform; this link only works for the first 100 clicks, then it’s kaput. I’ll eventually post the pre-published version of the article on this website, as I can freely share it. Much of the summary data from the research and the scripts I wrote for processing and analyzing the data are available on GitHub.

Background

The idea for the paper sprang from my work helping graduate students compare estimates over time when making maps in GIS. I wrote a post about these early experiences, where I described the fairly involved process we had to take to prepare and assess the estimates, only to discover that statistically they were not comparable (you couldn’t discern actual change from sample variability), or if they were comparable they were too imprecise to be useful. I wondered, how feasible is it to compare estimates over time – is going through all of these steps ever worthwhile?

To answer this question, I chose 25 detailed tables that captured a broad range of variables and represented basic, fundamental tables that many researchers would likely use. These tables included over 300 individual variables. For example, table B01001 included 49 variables: total population, total population that was male and that was female, and then 5-year age cohorts by sex (male population aged 0-4, male population aged 5-9, etc). I chose 19 geographic summary levels that represented the most essential geographies, from the nation down to census block groups. I compared estimates for these variables from the 2010-2014 to the 2015-2019 ACS for approximately 400,000 individual geographic areas; I chose these two 5-year periods as they do not overlap, and they are drawn from the same geographic vintage (2010) so I would not have to account for geographic boundary changes that occur with each decennial census.

ACS TABLES INCLUDED IN THE STUDY

Table IDTable NameVariablesObservations (M)% Total
B01001Sex by Age4919.918.0%
B01002Median Age by Sex31.21.1%
B02001Race104.13.7%
B03002Hispanic or Latino Origin by Race218.57.7%
B05002Place of Birth by Nativity and Citizenship*15†2.82.5%
B07204Geographical Mobility in the Past Year*19†3.53.2%
B08006Workers by Means of Transportation to Work*17†3.12.8%
B11001Household Type (Including Living Alone)93.73.3%
B13002Women Who Had a Birth in the Past 12 Months*193.53.2%
B14001School Enrollment by Level of School*101.81.7%
B15002Sex by Educational Attainment3514.212.9%
C17002Ratio of Income to Poverty Level83.22.9%
B19001Household Income (Dollars)176.96.2%
B19013Median Household Income (Dollars)10.40.4%
B19083Gini Index of Income Inequality*10.20.2%
B23025Employment Status72.82.6%
C24030Sex by Industry for Employed Civilians29†11.810.6%
B25002Occupancy Status31.21.1%
B25003Tenure31.21.1%
B25010Average Household Size of Occupied Housing Units31.21.1%
B25063Gross Rent24†9.78.8%
B25064Median Gross Rent (Dollars)10.40.4%
B25070Gross Rent as a Percentage of Household Income114.54.0%
B25077Median Home Value (Dollars)10.40.4%
B26001Group Quarters Population*10.20.2%

* Data for block groups and tribal block groups is not published for this table
† Value represents a subset of the total variables in the table; some variables were excluded

I opted against randomly sampling the variables, which would have allowed me to statistically analyze the results and make “firmer” assumptions. The primary audience for the paper is a practitioner-based one, for whom it’s more meaningful to have complete results for all geographies, for all or most variables in tables that people commonly use. The variables I chose also represent relatively broad swaths of the population. A random sample would have yielded many variables for small populations from disparate tables, which would be of limited interest.

I calculated change and percent change for each pair of estimates from the two time periods. I applied the test for significant difference, to determine if the estimates from both time periods were truly different from another, or if any difference was likely the result of sampling variation. If the estimates were significantly different, I calculated the coefficient of variation, to determine how reliable the estimates were. I summarized the results by variable, table, and geographic summary level. Ultimately, I had a dataset with 110.5 million individual data points that represented change for a variable between two time periods for a given place, of which 107.5 mil­lion were distinct (i.e. variables included in more than one table were only counted once).

Results

The results were surprising. Approximately 90% of the time, estimates from the two time periods were not statistically different from each other, and thus could not be reliably compared to compute change over time. The estimates that were significantly different were of low precision approximately 90% of the time (a CV value higher than 30). In short, it’s usually not worth investing the effort that’s involved with comparing two, consecutive 5-year period ACS estimates over time.

That being said, there are exceptions. Not surprisingly, the larger a population was, or the larger the sample size, the more likely it was that the estimates would be comparable and reliable. For the nation as a whole, estimates were comparable about 94% of the time, and of those that were comparable, estimates were of reliable precision about 98% of the time (a CV value less than 30). Both of these values fall to 73% for states, to 28% and 34% for counties, and 12% and 6% for census tracts. When considering all estimates regardless of geography, variables that represented large segments of the population were also more comparable and reliable. For example, estimates for the total population, median home value, population not enrolled in school, and workers 16 years and older were comparable 25% to 30% of the time and reliable 22% to 33% of the time.

There were some interesting divergences. Estimates for workers who take various forms of public transit to work ranked low in terms of being statistically comparable, but of the estimates that were comparable, they tended to be the most reliable. While this population is small or zero in many parts of the country, it is large and highly clustered in dense urban areas, so when the values are comparable they are of high precision. So if you were doing research in New York City or Chicago, you would have good estimates to work with (but not so if you were looking at all tracts / places / ZCTAs in New York State or Illinois).

Beyond the relationship between population and sample size, areas or populations that experienced a large degree of growth or decline were also more likely to be comparable and reliable, as these large shifts outweighed any sample noise. These two time periods experienced dramatic growth in home value (adjusted for inflation) and decline in unemployment. Median home value was consistently ranked as one of the most comparable and reliable estimates, while the unemployed labor force was ranked as one of the most comparable (despite the fact that this population is relatively small). The map in the header of this post depicts census tracts in and around San Antonio, TX. Rapidly growing exurban tracts tended to have a higher number of reliable estimates (CV values less than 30).

As a data librarian, I was equally interested in the means to the end, as a large part of my mission is to support students and faculty with this process. During the data collection and analysis phase, I wrote a post that provided tips on working with big(ger) datasets when your computing power is limited to your laptop. I relied on a PostgreSQL database to efficiently store the data and generate the calculations, and I had to be thoughtful in how to use Python so that I didn’t run out of memory, eschewing popular tools that make things easy in favor of approaches that were more efficient, and using SQL and Python in tandem to do the work that each was best suited for.

While the ACS isn’t well suited for short-term historical comparisons, it remains a valuable and indispensable dataset. Its strength is that it gives us a detailed snapshot of our present circumstances, at a fine degree of geographic and temporal resolution that is unmatched by any other US government dataset. We should not confuse shortcomings with the data with calls to curtail or eliminate it. The current administration’s attempts to sabotage the federal statistical system are detrimental to our country’s economic and social well being. As data practitioners we need to support and protect our nation’s datasets, while collaborating with agencies to improve them.

Library buffers and schools

Intro to GIS with QGIS Tutorial Updated

I posted the latest version of my Introduction to GIS with QGIS tutorial manual, updated for QGIS 3.44 Solothurn. The manual and sample data are freely available from my lab’s tutorials page.

This year’s changes are minor, as there weren’t major software updates that would impact the exercises. I updated the sample data as it was growing a bit stale, which necessitated updating screenshots and references in the text. I also updated the source LaTeX code to ensure the PDF meets the new federal WCAG guidelines for digital accessibility. I added a brief troubleshooting section on using QGIS with a Mac to the Introduction, as there have been a creeping number of annoying problems that have thrown my workshops off kilter (MacOS security blocking installation and writing of files to the user’s documents folder, and UI issues with the default color scheme and hiding menus in the background).

It’s hard to believe this is the 16th edition of this manual. When I wrote the first version back in 2011 (it was originally called Introduction to GIS Using Open Source Software), there was relatively little documentation on QGIS. In launching a workshop series and on-going support for it, I felt that I needed to create a basic guidebook. QGIS was far more primitive back then; converting the CRS for a GIS file required a separate command-line program, and you had to calculate natural breaks by hand! QGIS has certainly come a long way since then, and has been widely adopted. I’ve updated my materials as the software has evolved, but overall I’ve stuck to the same plan in terms of content, with modifications to the examples:

  1. Introduction: a general overview of the workbook, goals for the material and workshop, and updates to the manual since the previous version.
  2. An Overview of GIS: a short narrative that describes basic GIS concepts, and open source software.
  3. Exploring the Interface: explanation of the interface, adding vector data and viewing and selecting features, adding raster data, web base maps, and project files.
  4. Geographic Analysis: a cohesive case study that illustrates the process for doing an analysis while showcasing fundamental operations including: tables joins, plotting coordinate data, selecting, filtering, and deleting features, and geoprocessing tools like intersection and buffering. My original case study was locating a new comic book store in NYC, subsequently modified to locating a coffee shop, and more recently identifying public libraries in Rhode Island that met criteria for a grant to host an after school program.
  5. Thematic Mapping: exercise for understanding coordinate reference systems and how they function in QGIS, transforming systems, classifying data, and creating a final map layout. The original example was mapping healthcare sector employment by US state, but eventually switched to voter participation in federal elections.
  6. Data and Educational Resources: strategies and suggestions for finding GIS data, and recommendations for learning more through tutorials and workshops.
QGIS Interface
QGIS 3.44 Solothurn interface in Windows 11, used for the 16th edition of the manual in 2026

I’ve always used a follow-the-leader approach in teaching the workshop, where we stay together and the group does the material step by step. Everyone is a little fatigued by the end of the day, so for the thematic mapping piece we pivot to using the workbook; I demo the material and folks work on their own. I primarily wrote the book as a takeaway, so people could refer back to it after the session was over. It was also useful for self-directed learning, and the act of writing and updating it helps me internalize the material.

The last couple of years have been challenging though, and I have been thinking about what I could do differently. There have been a number of articles recently that discuss declining attention spans and students’ diminishing ability to thoughtfully engage with text (for example, in The Economist and The Atlantic). My own experiences bear this out. Increasingly, participants have a hard time following me in the session, and when I have student employees do the tutorial on their own as part of orientation, it takes them much longer and they struggle to finish. I have other workshops where I don’t use the follow-the leader approach, and instead I demo the material and everyone works on their own using the text, and once everyone finishes we move on to the next part. This works no better, as many participants wander away (mentally and physically) from the exercises and struggle to relate the text to what’s on the screen. Now, a good deal of success hinges on having good instructions, but I update them every year, and have used the same system and material in earlier years to good effect.

I have considered creating a more interactive web version of the workbook, or creating videos. The problem with both is that updating them each year would be far more time consuming; in contrast, updating the text and images in the workbook is pretty straightforward. I had one colleague suggest that my “old school” workbook fills an important niche in terms of structure and format, and is useful “as is” for deep learners. A professor I work with, who has had similar experiences in the classroom, expressed that there’s a limit to what an instructor can do; we can only simplify the material so much, and it’s up to students to rise to the occasion.

Reflecting on how I learn, I have gradually moved to videos to learn new material: whether it’s understanding a particular GIS method or tool, mastering a video game, or figuring out how to replace the belt on my clothes drier. Nothing beats a book for getting a comprehensive introduction to a topic, but the videos are helpful for targeted, stand-alone tasks. They can be comprehensive too, if they’re thoughtfully designed as part of a series. I’m definitely going to keep the workbook manual going, but am considering a video version for the future. I’ll wait until the 4.x version of QGIS becomes the new long term release (early next year) before I head in that direction, as the jump from 3.x to 4.x is more likely to come with some interface and functional changes.

The workbook has gotten a lot of miles these past 16 years, and I know others have used it for their own workshops and courses. Feel free to reach out if you have any suggestions.

Old QGIS Interface
QGIS 1.5 Tethys interface in Windows XP, used for the 1st edition of the manual in 2011
The Camino Primitivo with a Way Marker

Talks and Travels: Conferences and the Camino

I haven’t been keeping up with posting, as these past few months have been atypical. April was devoted to attending conferences and giving presentations. Much of this was prompted by my recent work with the Data Rescue Project, and the HIFLD Open rescue initiative in particular. The month began with a panel at Brown’s Data Science Institute, where the topic was Trust in Data. A few days later, I joined local colleagues at the Northeast Higher Ed GIS Facilitators Meet Up in Worcester, MA. I was honored to serve as the keynote speaker for the annual Big 10 GIS Conference (held virtually), where I presented on preserving federal datasets and the HIFLD Open rescue initiative. Shortly thereafter, I traveled to the Census Bureau’s headquarters just outside of DC for FedGeoDay 2026 and served on a panel of non-federal data providers who are contributing to the national data ecosystems. I came back to Providence just in time to give a poster presentation on our GIS and Data Services at the CHAIRS-C conference (Center on Heat, Health, and Aging Innovation and Research Solutions for Communities at the Brown University School of Public Health).

Then in May, I went off the grid. My wife and I traveled to Northwestern Spain to walk the Camino de Santiago, or Way of St. James. Established in the Middle Ages, the Camino is a series of routes that drew Christian pilgrims from throughout Europe to the Cathedral in Santiago de Compestella, which is believed to be the resting place of the apostle St James, brother of St John. We walked the Camino Primitivo, which is the “original” route established by King Alfonso II circa 814 AD. It’s also considered to be the most challenging of the routes, as it climbs through mountains and forests before descending into farmland. It’s not considered a “wilderness” hike however, as all of the routes follow a mix of unpaved and paved roads between towns and villages. A map of the primary routes (from Wikipedia) is below. The French Way is considered the primary route and is the most heavily traveled. The Portuguese and Northern Ways are also popular, followed by the Primitivo.

Map of the Primary Camino Routes
By WikiPate – credit to Sémhur, and for the logo: Manfred Zentgraf, CC BY-SA 4.0, Map from Wikimedia

The paths are marked at regular intervals, and whenever you have to turn or change direction. You look for a white or grey stone marker with a scallop shell, a symbol of St. James and of the Camino, to guide you. In places where a marker stone isn’t feasible, a blue and yellow tile of the shell is embedded in a wall or building to point the way. The system was so good that we rarely needed our phones; we used nothing more than our eyes and an excellent guidebook with detailed topographic maps of each stage of the journey, elevation diagrams, and a directory of landmarks and places to stay (the Village to Village Camino Guides – highly recommended).

The routes run into the hundreds of kilometers; the Primitivo is a shorter trek that covers about 320 km between Oviedo and Santiago, but in exchange for the shorter distances you have greater changes in elevation. As this was our first attempt at something like this, we opted to do half the route, beginning in a town called Grandas de Salime, located at a large dam and reservoir as you leave the region of Asturias and enter Galicia.

Embalse de Salime
Our starting point: view of the Embalse de Salime from the Hotel Las Grandas

Accommodations and cafes serve pilgrims throughout the route; albergues offer a mix of hostel-like rooms (bunk beds in shared rooms) and single rooms, and there are also basic hotels. Our 183 km walk took us 9 days (3 days walking / 1 rest day in Lugo / 5 days walking). You are issued a pilgrim’s credential or passport when you begin, which grants you access to pilgrim-reserved accommodations and resources. On your journey, you need to get your passport stamped twice a day to verify that you are doing the walk. You always get a stamp at the places you stay, and in-between you can pick up others at cafes and restaurants, churches, museums, visitor centers, and even certain stores (we managed to get one at a cheese shop). Once you reach the cathedral in Santiago, you visit the pilgrim’s office (essentially the Camino DMV), where you present your passport to receive the Compostella, the official document that certifies that you finished the pilgrimage. You need to walk 100km minimum (200km if you’re biking) to qualify.

Camino Passport Stamps
A Selection of Stamps from my Pilgrim’s Credential or Passport

It was a deeply moving experience, retracing the steps that countless pilgrims took over a thousand years, and ending in front of the statue and tomb of St James behind the high altar in the cathedral, receiving the Eucharist at the Pilgrim’s mass. It was a relief to disconnect from technology and work, boiling life down to the singular goal of getting from point A to B each day. It was physically satisfying, pushing my body to walk 10 to 20 miles a day in rough terrain in all kinds of weather. It was wonderful to meet new friends; there is a cohort of people who happen to begin their journey simultaneously with you, and you see them throughout the walk, sharing the road for a time or a meal at the end of the day at the albergue. And it was a lot of fun, for a geographer who enjoys navigating a landscape with no digital do-dads, and who loves collecting stamps!

An experience like this alters your perspective, and it’s been difficult to transition back to my normal routines. It has strengthened my belief that it’s time for me to consider new possibilities and next stages in my career. Please reach out (via LinkedIn or email in the sidebar) to share opportunities. My resume is available on the About page.

Stay tuned for some heat and climate-related dataset suggestions in my next post; resources I compiled for the heat conference, and new ones I’ve learned about at FedGeoDay.

HIFLD Next Map Preview

HIFLD Next GIS Data Catalog

Last December, the Data Rescue Project (DRP) finished an initiative to download and archive over 400 geospatial data layers from the defunct HIFLD Open repository to DataLumos, an ICPSR-sponsored repository for federal government datasets. I wrote this brief post that summarized our work.

The Public Environmental Data Partners and Fulton Ring have launched a new community-shaped hub for finding, previewing and downloading GIS data collections, and its debut HIFLD Next collection is built on this rescued dataset. The portal increases the accessibility of the HIFLD Open data, with enhanced options for searching, previewing attribute tables and layers, and downloading or streaming data in a number of different formats. Beyond publishing a website, the group is hoping to build a community of practitioners around this project to support and sustain it, and to provide updated datasets and additional collections in the future.

They are keen to solicit feedback from GIS data users, and particularly from librarians and data specialists who provide active user support and who would potentially refer to the portal as a source. After you’ve explored the portal, feel free to submit feedback via their survey.

To learn more about the project, you can read this press release from the PEDP and this announcement from the DRP. The project is likely to be a primary topic of discussion at FedGeoDay 2026, which takes place in late April in Washington DC.

HIFLD Open DataLumos Archive

HIFLD Open Data Archived in DataLumos

Some good news to end the year: the Data Rescue Project has finished archiving all of the GIS data layers that were in the HIFLD Open portal, which was decommissioned at the end of summer. I wrote a post for the DRP that summarized the work we did, and you can find all the layers in ICPSR’s DataLumos repository, where you can search for and download layers one by one. I also archived the index for the series and a crosswalk that DHS published for locating updated versions of the data from the individual federal agencies that created them. If you wanted to download the entire set in bulk, it can be transferred from the Brown University Library’s GLOBUS endpoint; there are instructions for doing this on our library’s Data Rescue GitHub repo.

This project was an archival one, in that we were taking a final snapshot of what was in the repository before it went offline. In the coming year, I’ll be thinking about approaches for consistently capturing updates, and there are some folks who are interested in creating a community-driven portal to replace the defunct government site. Stay tuned!

2025 has been a tough year. Wishing you all the best for the year to come. – Frank

DataLumos HIFLD Open Archive

HIFLD Open Shutting Down

HIFLD Open GIS Portal Shuts Down Aug 26 2025

HIFLD Open, a key repository for accessing US GIS datasets on infrastructure, is shutting down on August 26, 2025. This is a revision from a previous announcement, which said that it would be live until at least Sept 30. The portal provided national layers for schools, power lines, flood plains, and more from one convenient location. DHS provides no sensible explanation for dismantling it, other than saying that hosting the site is no longer a priority for their mission (here’s a copy of an official announcement). In other words, “Public domain data for community preparedness, resiliency, research, and more” is no longer a DHS priority.

The 300 plus datasets in Open HIFLD are largely created and hosted by other agencies, and Open HIFLD was aggregating different feeds into one portal. So, much of the data will still be accessible from the original sources. It will just be harder to find.

DHS has published a crosswalk with links to alternative portals and the source feeds for each dataset, so you can access most of the data once Open HIFLD goes offline. I’ve saved a copy here, in case it also disappears. Most of these sources use ESRI REST APIs. Using ArcGIS Online or Pro, and even QGIS (for example), you can connect to these feeds, get a listing in your contents pane, and drag and drop layers into a project (many of the layers are also available via ArcGIS Online or the Living Atlas if you’re using Arc). Once you’ve added a layer to a project, you can export and save local copies.

QGIS ESRI Rest Services
Adding ArcGIS Rest Server for US Army Corps of Engineers Data in QGIS

If you want to download copies directly from Open HIFLD before it vanishes on Aug 26, I’ve created this spreadsheet with direct links to download pages, and to metadata records when available (some datasets don’t have metadata, and the links will bring you to an empty placeholder). Some datasets have multiple layers, and you’ll need to click on each one in a list to get to it’s download page. In some cases there won’t be a direct download link, and you’ll need to go to the source (a useful exercise, as you’ll need to remember where it is in the future). Alternatively, you can connect to the REST server (before Aug 26, 2025) in QGIS or ArcGIS, drag and drop the layers you want, and then export:

https://services1.arcgis.com/Hp6G80Pky0om7QvQ/ArcGIS/rest/services

I’m coordinating with the Data Rescue Project, and we’re working on downloading copies of everything on Open HIFLD and hosting it elsewhere. I’ll provide an update once this work is complete. Even though most of these datasets will still be available from the original sources, better safe than sorry. There’s no telling what could disappear tomorrow.

The secure HIFLD site for registered users will remain available, but many of the open layers aren’t being migrated there (see the crosswalk for details). The secure site is available to DHS partners, and there are restrictions on who can get an account. It’s not exactly clear what they are, but it seems unlikely that most Open users will be eligible: “These instructions [for accessing a secure account] are for non-DHS users who support a homeland security or homeland defense mission AND whose role requires access to the Geospatial Information Infrastructure (GII) and/or any geospatial dashboards, data, or tools housed on the GII…“