First Project List
List Of Acronyms
- CAD
- Computer Assisted Dispatch
- NENA
- National Emergency Number Association
Introduction
I’ve been silent lately because my wife and I have been busy launching our new LLC, Dunsworth, Mann, & Associates, LLC. Part of that process, aside from an interminable amount of paperwork, is to start putting together projects and products that we can offer to clients. The first of these is something that I’ve been working on for a while now, and that is a synthetic data generator for 9-1-1 call centers. The idea is to create a tool that can generate realistic, but entirely synthetic, data that can be used for training, testing, and research purposes without compromising the privacy of real individuals.
The tool is named SynthCCD. While it is focused on the creation of synthetic data for the 9-1-1 industry, we also believe that it can be used in various other industries with some customization. My intention is to have a product that we can license to educators, companies in the 9-1-1 sector, analysts, and other consultants who would like to work on various projects of their own, but eihter do not want to use or cannot use data from a 9-1-1 centre. I will discuss the possible use cases as I discuss some of the design desicions. The link to the software’s page will be at the bottom of the post.
Now, before many of my friends say anything, yes, I’m selling licenses to use the software. However, when you buy a license to use it, you will get the source code, so it’s open-sourced, but it won’t be freeware. Yes, I would like other people to contribute back improvements they may make to the code over time. However, the truth will still be this; I might request it, but I don’t know that I can enforce it.
Back Story
In 2024, I was preparing to give a presentation with a friend of mine at the 2024 NENA Annual Convention in Orlando, FL. Part of what I wanted to do with this was give a live coding demonstration to show, in real time, how a request from operations could be processed through analytics, and return a real business impact. Essentially, I wanted to show how data science was relevant to 9-1-1 centres around the country and emergency communication centres around the world. My concern was balancing showing that with exposing real data from a real centre. Thankfully, I work in a supportive centre and as long as I cleaned the data to ensure that no personal information was ever used, I was allowed to use my home centre’s data, this time.
I knew that I couldn’t count on that in the future, so that started the germination of what is now becoming SynthCCD. At first, I only needed something that would emulate incident data that you could find generated from any Computer Assisted Dispatch (CAD) system. So I tried Mostly AI and the results weren’t bad except that I had to give them real data to get good results. They offered me time and processing power to build sets and see what I could do with their project. While the results weren’t bad, they weren’t what I was hoping for. So, I started thinking about how I could do it myself.
I knew that I could leverage the faker library to build something myself. The first iteration of this is on GitHub. It worked, but I found out that it had a few problems with it. When I had created a dataset for a presentation at the Virginia Conference on Innovation in Public Safety at Randolph-Macon University in Ashland in 2025 I found that all of the elapsed times between events were normally distributed, as were the volumes of events created per hour and per day. While the presentation went well, I was very frustrated that my data didn’t look and act real. That led me to both version 2 and 3 of this project. I found that I could work with NumPy and use different functions to create the proper distributions. The next challenge was identifying the right distributions with the right parameters. I worked with different AI-based coding assistants to help me identify the correct distribution with the right parameters. It’s a lot easier that way. Once I had those identified, I could build everything out. Yes, for those asking, I did use AI-assisted coding harnesses to speed up the dev work. That’s a decision that I don’t take lightly, but it has sped up my work considerably.
Version 3
So, with all that in the background, the idea was to take the pieces of the first two versions that I thought worked and start building. I started with a SQL query that I used in my day job. That allowed me to start identifying the fields that I wanted in the incidents data when it was created. After I added to the base schema, I started creating a new TODO list to add those elements into the system.
The first thing that I worked on was getting the distributions even more accurately defined. The more that I worked on this, the more that I wanted to ensure that the output looked more and more like real data. I had even started finding more data from data.gov to ensure that I had distributions of categorical data as correct as I could. That started making the output data feel a little better. However, the next challenge I found was that I had the right city name, but the addresses were all fake, literally, and it was a challenge to have a diverse enough address set to pass muster. So, I figured out how to leverage Open Street Map to generate real addresses for real locations. The next step became more in the line of vanity than what I really needed. I tried to work out how to scale the number of events of generated weekly using the population. I also extended that to the volume of calls to a PSAP. Being honest, the latter was the reason I was working this out. I could make up numbers for Alexandria because I know that data back and front. However, if I move up and down in population, I am not so certain. After I had these down and I could prove the scale, I felt it was ready to write up and make public.
Other steps
In the development process, I had to decide how the user is going to interact with the software. I had considered building a GUI, but in the end, I decided against it and only made command-line and TUI options. I asume that most users will likely interact with it for data at scale. I did set up that a user can write up a JSON, YAML, or TOML file with the parameters to be passed in the command line. All of the parameters can eqaully be defined in the TUI. I also updated the security so I can reduce the exposure as much as possible. I’ve also set up the program to output the data in CSVs, parquet files, pandas and polars data frames, and now output into various databases. I know that is a lot of input and output options, but it allows for different user groups to use this software as they need.
What’s coming next?
I have a couple of different ideas for what comes next. The first thing I’m thinking of is the ability to create synthetic Emergency Incident Data Objects (EIDOs) to test connections, data transfers, and the ability to successfully process them. I’m also looking at building an AI agent companion to assist someone in creating everything and have that be the equivalent of a GUI. I think that these are good enhancements for the future.
What do you think?
So, as per the question starting this section, after you’ve looked at it and read the extensive documentaiton, what do you think I should be doing next with this as it grows?