Benji to Penji
Before Penji there was Benji, the first system I ever built for anyone. It ran the data for a national diabetes prevention program across 11 organizations in 5 countries and U.S. territories, and it showed me what Penji had to be.
Benji: 2018 until the program's CDC funding closed · Penji: 2019 to now
The program
In 2018 I joined the Association of Asian Pacific Community Health Organizations (AAPCHO) part time, doing data entry for the Pacific Islander Diabetes Prevention Program. It brought the CDC National Diabetes Prevention Program to 11 organizations across the U.S., Palau, the Federated States of Micronesia, the Marshall Islands and the Northern Mariana Islands.
The 11 organizations were the delivery sites. Each ran its own program with its own program coordinators, data specialists and lifestyle change coaches, in offices that were sometimes trailers, where power, internet and even the class venue changed from week to week, and many of the staff were volunteers. Our central team of two or three supported all 11 sites, and I managed every site's data closely. The program had to prove with data that it met the CDC Diabetes Prevention Recognition Program (DPRP) standards, and at the start that data was on paper.
When I started, only about one in five participants' records could be submitted to CDC.
What I built
There was no data system when I joined. From the data entry work I saw what the program needed, took it on myself, and built my position around those duties. I first tried a React web app, but in 2018 a secure web application was beyond what I could build, so I built Benji from scratch in Google Sheets and Apps Script.
- Secured Google Drive folders, one per organization, so each organization worked only with its own data.
- Separate sheets that worked as linked tables, a kind of relational database built on spreadsheets: coaches with their coach IDs, cohorts with a cohort ID and the coach ID of the coach who led it, and registrations with each participant ID. Linked by those IDs, the sheets came together into one system I could query.
- Session logs that staff filled in from the paper sign-in sheets after each class. Each weight was compared with the participant's previous weight, and a change past set thresholds turned the cell yellow, orange or red, so the person entering it could check it with the paper still in front of them.
How the data moved
- Organization foldersOne secured folder per organization, holding its classes and session logs.
- Log GathererCrawled every organization's folder and pulled every class's session logs into one place, more than was needed, because it couldn't tell which logs were new.
- MergerCompared what the Gatherer collected against the logs already in the Processor, found new rows and edits to rows already collected, and merged only those changes in.
- CleanerRan error finders and outlier checks across the combined data. It began as a clone of the Processor, split off to share the load.
- ProcessorRead each participant ID, found when the participant and their cohort started, worked out which six-month DPRP submission each record fell into, projected whether and when each organization would earn recognition ahead of CDC's own evaluation, and tallied each site's results against milestones modeled on the Medicare Diabetes Prevention Program billing codes for the program's pay-for-performance payments.
What I found in the field
I read the data and visited the sites, starting in Chuuk, then Kosrae in the Federated States of Micronesia and Ebeye in the Marshall Islands. Most of the errors traced back to how the data was collected.
- Weighing
- People were weighed on different scales, on different surfaces, read from different angles. I pushed for funding so every site had enough scales for each participant to be weighed on the same scale every class. A scale that reads a few pounds off is fine if it reads the same way every week, because the program measures percent weight change.
- Participant IDs
- Staff were numbering participants in the order they signed in that day, so the same person could be number 1 one week and number 4 the next. I trained staff that a participant ID belongs to one person for the whole program.
- Entry mistakes
- Typos and misread weights were caught by the color flags in the session logs at the moment of entry.
The training and the system changed together: each round of training came with an update to Benji, and each update came with training.
Results
- 96%average share of participant data CDC could accept, within a year, up from about one in five; some organizations reached 100%
- 4,500+participants over the program, about 750 a year
- 8 of 11delivery sites reached CDC Full Recognition, and so did AAPCHO
- $1M+in pay-for-performance payments to the sites over six years, by my count
The 96% held through the end of the six-year project. AAPCHO took ownership of Benji in 2019.
Then Penji
Benji worked, and it reached the limits of spreadsheets: sheets with hundreds of columns slowed as the data grew, and checks ran after the data was in.
Benji ran until PI-DPP's CDC funding under DP17-1705 closed, which ended the program's activity at those organizations. A couple of years later the work relaunched under a new sustainability plan, and that relaunch is what Penji has been built for.
Benji's web rebuild began in late 2019, in React and then in Angular from 2020, and became Penji, the platform I have built as my own project since. Penji keeps what Benji got right: each organization works only with its own data, the program's standards are checked as the data is entered, and it works on connections as slow as 28 kbps that drop in and out. Pasifika Health Data Center, built for AAPCHO, is its first product, and Benji's users are moving onto it.
Four working demonstrations rebuild parts of Penji on synthetic data:
The code is on GitHub at github.com/seanpatrickrodriguez/penji-demos.