Podcast: Play in new window | Download
Subscribe: Apple Podcasts | Spotify | Amazon Music | Android | Youtube Music | RSS

Brad Bolliger entered the knowledge graph space via enterprise software system design and data analytics. That background informs their pragmatic and strategic approach to the use of semantic technology in systems that facilitate information exchange across government agencies.
We talked about:
- their work at EY (Ernst & Young) on data and analytics strategy assessments and enterprise software design and as a co-chair of the NIEMOpen Technical Architecture Committee
- how their work on EY’s Unified Justice Platform introduced them to the knowledge graph world
- a quick overview of entity resolution
- the NIEM standard, its origin in the wake of 9/11, its scope, how it’s built and managed, and how governments use it
- their pragmatic approach to ontology and vocabulary management
- the benefits of the extensibility of the RDF format and knowledge graph technology
- how entity-centric data modeling accelerates and facilitates systems evolution
- their take on “analytics enablement engineering”
- their approach to crafting AI-ready data and building AI-aware enterprise solutions
- some of the neuro-symbolic AI architecture’s they have seen and implemented
- their call for more systems thinking and systems analysis to create more effective services that work together in a more ethical and effective way
Brad’s bio
Bradley Bolliger (they/them) works in the AI & Data practice of Ernst & Young and serves as co-chair of the NIEMOpen Technical Architecture Committee, an OASIS open standards project for data interoperability.
Brad assists clients across various industries with optimizing data platform ecosystems, enhancing customer relationships, and leveraging advanced analytics tools and techniques in their digital transformation efforts. In addition to designing data platforms and AI/NLP systems, Brad has served in lead analyst roles for public sector information system modernization efforts, including major contact center data ecosystems and integrated criminal justice system environments, the latter of which would lead to the development of the UnifiedJusticePlatform.
Connect with Brad online
Video
Here’s the video version of our conversation:
Podcast intro transcript
This is the Knowledge Graph Insights podcast, episode number 42. When you have to account for the people and other entities involved in high-stakes situations, you need a system that delivers accurate, unambiguous information. Brad Bolliger does this in their work on EY’s Unified Justice Platform. Brad is relatively new to the graph world and has adopted a pragmatic approach to semantic modeling and knowledge graphs, focusing on applying lessons learned in their extensive experience in enterprise systems design and data analytics.
Interview transcript
Larry:
Hi, everyone. Welcome to episode number 42 of the Knowledge Graph Insights podcast. I am really delighted today to welcome to the show Brad Bolliger. Brad works in the AI and data practice at EY, the big consultancy in Chicago, and also helps co-chair the NIEM Information Exchange, the Info Exchange Network and standard. Welcome, Brad. Tell the folks a little bit more about what you’re up to these days.
Brad:
Thanks for having me, Larry. I’m thrilled to be talking to you today. Yeah, I’m non-binary. I use they/them pronouns, and I work in the AI and data practice at Ernst & Young, as you said, where I do data and analytics strategy assessments and enterprise software design, things like that. I’m also co-chair of the NIEMOpen Technical Architecture Committee, which is an Oasis Open standard for sharing data in public services primarily, but for specification for developing information exchanges. And I’m working on semantics and software design more generally.
Larry:
Yeah. And you kind of not stumbled, but you had semantics thrust upon you in this new role, I understand, ’cause one of the projects you work on, I don’t know if you’re still working on it, was the Unified Justice Platform at EY. Can you talk a little bit about that and how it brought you into the semantics world?
Brad:
Yeah, that’s right. It spun out of an assessment from a county government wanting to overhaul their integrated justice system, which was the collection of actors who collaborate or have this adversarial relationship to administer the process of justice in their jurisdiction. And because very often they’re their own elected officials with their own budgets, they have their own software to fulfill their own functions. And that means that they are kind of inherently operating a distributed system, sending messages back and forth to say, “Hey, we booked this person into the jail. Hey, we’ve got this court date coming up. Hey, we’re filing these charges.” And they need to orchestrate complex operational processes across multiple software systems and multiple groups of people, again, kind of across jurisdictions or enclaves. And that was, of course, a really interesting systems analysis process that led to the development of a solution to this problem we were trying to assess, which we later called the Unified Justice Platform and is an event-driven architecture for building an entity-resolved knowledge graph as an operational data store programmatically as messages are exchanged between the stakeholders in the Enclave.
Larry:
Yeah. And you used a couple of words in there. I want to clarify for folks who might be new to them. The notion of entity resolution, the entity-resolved knowledge graph, I’ll just point out that we met through our mutual friend, Paco Nathan, who works for Senzing, a company that just does entity resolution. And can you talk a little bit about entity resolution, how that fits into the needs of this distributed system and how you implement it in the platform?
Brad:
Yeah. Actually, I’ll plug almost two years ago, we did a webinar with someone from Senzing and talked about the fundamental utility of entity resolution and relevance, I suppose, as a problem more generally. Entity resolution is essentially about creating, for me, is essentially about creating a high quality master index of whatever kind of data that it is that you’re looking at. So in this case, we were talking about a master person index so that you have a more reliable picture of the same natural person, no matter which software system is representing the data that describes the person subject to judicial proceedings in particular. But thinking about entity-centric data modeling more generally, you got a different type of entity, you still need to disambiguate which location you’re talking about, which person you’re talking about, which entity that really is. And if there are different representations, different records that relate to the same underlying entity, that process of entity resolution therefore has this really broad systemic benefit to data management and data engineering in particular, because ultimately it’s about the master index at the end of the day.
Larry:
Yeah. And as you talked about that, you mentioned that it’s like this a canonical record of entities. And how does NIEM fit into that? Because that’s a vocabulary as I understand it.
Brad:
That’s right.
Larry:
Yeah. Can you talk a little bit about NIEM and how that works with entity resolution?
Brad:
Yeah, very briefly on NIEM, NIEM spun out of the post September 11th realization that public services needed to share data to collaborate more effectively to actually solve emergencies, but just problems in general. And what they realized was that they need to have a common language to collaborate more effectively. Again, because systems, machines, software systems, have this really concrete definition of we use these particular terms and they mean something in our enclave, but you could have a person’s full name and a person’s first name and a person’s last name in two different records, but actually they’re the same real person. So NIEM came out of an attempt to at least address some of that disambiguity. And what is most interesting to me about NIEM, honestly, is that it is a collaboratively defined list of vocabulary. So we actually get domain participants involved and they decide we use these terms and they mean these things.
Brad:
And so it’s an attempt to reduce the amount of complexity that you could use to describe a different person, but communicate the same meaning without losing the information that’s entailed in some data record. But I’m digressing a little bit probably. What NIEM is a framework for building message specifications, APIs, if you like, or other types of structures, data structures in general that is a community agreed-upon set of terms that have some kind of core relevance, person, entity, organization, or have some domain specific function, like, subject or something in human services and so on.
Larry:
Interesting. Yeah. And as you talk about that, that attempt to align people on vocabulary is such a notoriously difficult problem. And I don’t know how many jurisdictions we’re talking about here, but every little town in America has a police department and other social services that they do. What is the scope or the scale of that? And is it facilitated in any way by existing standards or vocabularies?
Brad:
Oh, very much so. In fact, the problem is even worse than you’ve described it very charitably, I think. Just in the United States alone, I’m told that there are over 18,000 law enforcement agencies, just law enforcement agencies. Nevermind how … Anyway, so NIEM is a voluntary open standard. So it is something that is available, but is usually not mandated. There are some places where it is mandated for specific types of services. So the scale of the problem that we’re talking about really depends on who’s included in the conversation. And that is one of the major goals of defining a NIEM domain is to bring as many of the participants in a community of interest to the table to align these different representations, again, for addressing some specific mission or function maybe more generally. The scope of it though is that it’s extensively used in many, but of course not all public services in the United States.
Brad:
And that’s at the federal, state, local, tribal level and interactions between some of those levels or types of organizations. So there will be regional collectives, for example, like, a joint power authority may have some regional thing or multiple states may subscribe to the same thing that’s not technically owned by the federal government, the states own it, like NLETS, for example, but it’s used extensively in the public services of several sovereign states in the world. It’s, I’m told how EUROPOL and INTERPOL share data, things like that.
Larry:
Interesting. And as you talk about that, and once you’ve done that hard work, and that could be a whole episode right there, I’m sure just talking about how NIEM can align on that terminology, but once you’ve got that done, then what do you do with it? I know you have a knowledge graph and you’re using it there. And so that’s like the use in the Unified Justice Platform that you do. Does NIEM, does that vocabulary serve other platforms and services as well? Or …
Brad:
Absolutely. It’s in use actively as the interface between a lot of different law enforcement agencies, justice agencies, all kinds of places, much more than justice and public safety though. So it’s in use in these software systems, collaborative information exchanges or collaborative analytics environments more generally. But what you do with it, for example, is, okay, you used the term canonical earlier to describe the resolved entity, but you take this core list of terms that are typically used in certain contexts, but you apply them in your own way. You’ve got some ontology that represents your opinion of what this combination of data structures means in your context. And so you localize those terms to your particular use case. So just in the Unified Justice platform, it is a template, so to speak, of backend facilitating infrastructure that you can plug in, but your particular producers and consumers will have their own extensions or customizations or we like to represent it this way and have this particular label at the header and not just contained in the depth.
Brad:
So okay, great. Customize it how you need and you hope to customize it as little as possible, but there are other requirements when you’re building production software and it’s not always about the semantics.
Larry:
Yeah. And so I mean, going back to the start, you’re like a systems architect and data person as well.
Brad:
That’s right.
Larry:
But one of the things, ’cause I love that you come out of that world and have come into this world. And one of the things that you’ve discovered as you’ve entered this semantic world is that it doesn’t always need full-blown, high-powered ontological stuff and just in full force description logic and everything. Can you talk a little bit about your philosophy around that kind of pragmatic approach you take to ontology and vocabulary management?
Brad:
Yeah. Well, I try to think about creative problem solving at the end of the day. What is the problem actually and what do we need to do to solve it? What do we have and what do we need? And I think the gist is, well, probably I should say to the semantic folks with all possible respect, I probably just have not been in a situation that required that full breadth of the semantic implementation in that way because I think what I’ve landed on with some of the folks that we know and have been talking to over the past couple of years in particular about this, is that that’s often, I was going to say a distraction.
Brad:
That’s not always as helpful to solving the actual problem as we might like to think. When we’re building production software, when one is building production software, you need to be clear about what you mean and reduce the complexity and variety as much as possible, but that’s kind of inherent to the way that systems evolve, networks evolve, that you also are to some degree trying to facilitate the proliferation of complexity at the ends of the network.
Brad:
You’re trying to make it easier for someone to add another software system or to change from one software system to another system, but not have the data break. And usually that means changing it less than you would like and not totally remastering it in a fundamentally different way ’cause that requires all of your other existing things to change as well. And so again, it is about organically evolving your ecosystem of software producers and consumers, and usually that means meeting your current production instances a little closer to where they are than where you would like to go in the future.
Larry:
Yeah, no, and that’s, again, the pragmatism of that. It also points, I think, to just the general benefit of the extensibility of the RDF format and knowledge graph technology in general. Is that a benefit you’ve discovered as you’ve got into this?
Brad:
Absolutely. The graph-based structure is really critical to enabling what I called the entity-centric data modeling or the person-centric data model in the case of the Unified Justice Platform, where we’ve got a couple of different software systems that all are saying, have different data about the same natural person, but even however many systems there are, each of those has a person record, because the software system itself needs to know which person we’re talking about so they can tell the other software systems and resolving that to the same kind of natural person is ultimately the goal of, well, an information exchange data model like this one, but the entity-centric data modeling is the kind of unique benefit of the graph to allow complexity again from different viewpoints on the same thing. They have different properties that they want to describe about the same person or the same thing or vehicle or location.
Brad:
And you want to make it easier to change the way that the software produces the data about that underlying thing. And there’s some really tactical examples of this that we’ve given on a couple webinars with NIEM in particular, is that if I have one of the participants in my exchange ecosystem, so one of the software systems is being replaced, upgraded to a new system, the data structure that they store the data in is going to be different, and therefore the way that they produce, publish the data is going to be different as well. And so if you have a person record relating to a court case record, this is a classic example. You can also have the kind of … Ultimately, what I’m talking about is enabling operational flexibility and separating that from the concerns of the way the data is managed. So I’ll get back to this person’s case record connection in a second.
Brad:
So you can allow for the structure itself to change by having this entity-centric picture, and that’s what’s so special about graphs is to just have a slightly different shape, but still maintain the connection between a natural person and the data property that you are looking for in particular for your given use case. And so I gave this example of there was a court system that was changing their court case management system to a new system, and that of course changes the data structure, but that happens all the time. You’re trying to make that easier. So a more specific example, this didn’t coincide with a change to a system itself, but they changed the data structure. So, for example, the court would book co-defendants on a case with one case record. There would be one case record with one or two or more co-defendants on that case for the same charges.
Brad:
And they found just as a matter of administrative convenience that it was actually easier to have one case record for a given person, so that each person gets their own case record, but still handle the court hearings in the way that you want, hear them at the same time or what have you, so that the judiciary has the flexibility and the court administrators have the flexibility to administer the process of justice separate from the data management. And what that changes to be really direct is, you now have one natural person with one case record, and instead of having one case record with two people linking to it, now you have the same person record pointing to one case each, super simple. So yes, you’ve got to change the software that works with that data, but you can make that simpler by just saying which person is connected to which cases, and it doesn’t matter if there’s one to end.
Larry:
As you talk about that, I’m reminded of … We talked earlier about in your role as a data analyst, but you’re also a data engineer. And so that you’re on both the consumption and the creation, so you’re equally cognizant and capable at both the consumption and creation sides of that, which also kind of leads into the notion of a data product. And what you were just describing that is like that operational stuff that any core, one core or something is doing, that’s just like one application of the use of this data. Can you talk a little bit about that separation of concerns between the application layer, the deep engineering layer, and how things like data, whether it’s a data product specifically or like that conception you were just describing of data doesn’t change, operations change.
Brad:
Yeah. Yeah. So this is what I liked about what we called analytics enablement engineering, which is data engineering to make analytics more useful or easier or enriched with new information, which is to say that they’re not kind of two ends of a linear process, that they are part of the same continuous process, honestly, that they’re kind of building blocks to solve problems at greater scale, honestly, is a big part of it. So when you need more meaningful signals or higher quality information, you want to change the system that produces that data. And very often as an analyst, you’re aggregating multiple things, doing some combination or correlations or statistical analysis of some kind, or if you’re doing some advanced analytics, some data science use cases on your underlying data, very often I found working as a kind of production data analyst for many years that you really got to go to the point of the source.
Brad:
And that’s not the production of the data product from a software system to across a silo to some other software system, but it’s at the point of creation. And this is of course not as literally true as it used to be, but in general, people create data, not machines. And that means that what the data is, where it comes from, I mean, especially when you’re talking about a case management system or a lot of e-commerce data, the data that you capture is a function of the form that you put in front of the user. And that’s how a lot of enterprise software works actually, is that there are forms that create structured data, which are usually just a digital representation of something that used to be paper with all due respect, but that’s really the kind of crux of what we’re talking about is that going to the source, which is the creation of the data and as your problem gets more complex, you’ve got to look at more sources, some of which are machine representations of events that occur, and therefore things can get really complicated really quickly and scale very quickly.
Brad:
But ultimately it’s about the data that people create and the data that people want to consume and not seeing the ecosystem as a black box or as a series of silos that one needs to navigate through, but seeing them as part of solving some problem that human beings have actually.
Larry:
Yeah. And I love the focus on the end human user and the fact that almost all this data comes from some human activity or entry or observation or something like that. When you talked a minute ago, you used the term analytics enablement to describe that. Another common application nowadays is AI, obviously it’s just everywhere. And one of the things you’ve talked about is AI enablement. What does that look like in the context of everything you’ve just said, how does the current … And AI is obviously a broad term that means different things in different contexts, but how in general would you say that, and I’m thinking particularly of the modern LLMs and that kind of stuff, how is that entering into the workflows and the management and the architectures that you’re working with?
Brad:
Well, that’s an interesting question. I’ve really enjoyed, it’s the perspective of my practice in AI and data at Ernst & Young that we call AI-ready data, and that it really is about not just mastering your data, but understanding it, preparing it and making it ready for AIs to use, but also for really understanding what you have so that you understand how the AI needs to be built and what it actually can do reliably. There are a lot of numbers floating around about pilots or language model-based chatbot use cases and the utility or risks that they may pose. But I think if you kind of zoom out longer term, you’ll see kind of fundamentally that it really depends on … What do I mean? It really depends on the nature of the problem that you’re trying to solve, I think that’s what I’m trying to say.
Brad:
And that a lot of folks are pointing out that, even using knowledge graphs for retrieval augmented generation doesn’t always solve the problem. And very often, and this is why I keep talking about software and data mastery and entity resolution and aligning semantics is that usually you don’t want to introduce probabilistic outcomes into a process. Very often, actually determinism is the thing that you want and you want some degree of creative flexibility, which is why I think you and I understand that AI is much more than just the transformer architectures. AI has been in place for a long time and it means something more fundamental than asking a pool of language correlations what something means. That’s not actually how you get the outcome that the user is looking for. And sometimes the user really is looking for that or is looking to collaborate and extrapolate on something.
Brad:
But in an enterprise context, you either can’t risk or don’t want that degree of ambiguity and uncertainty in the outcome when really what you’re looking for is a slightly fuzzy way to repeatedly solve a problem. And so that really requires you to understand your data and understand what you can do with it. And I think that’s what’s kind of different about my perspective relative to some of the others is that it’s not just about what a language model kind of computes. There is some underlying thing that’s happening, and very often a different approach is a better, cheaper, faster solution.
Larry:
Yeah. I think we’re already starting to see those different approaches. At least I’m inferring that from the results I’m getting from some of the AI applications out there. But you’re reminding me too of another thing, I don’t know if you saw Tony Seals talk at Connect to Data London, or if you’ve seen his posts about the neuro-symbolic loop and that natural and obvious complimentary nature of the neural network-based learning stuff and our work, the stuff I think of is my work, the more symbolic side of the equation goes back … And that’s one of the oldest things. That goes back to the dawn of AI back in the ’50s and ’60s when they’re talking about connectionists versus, what do they call it, connectionist versus symbolic, I think was the distinction back then. The way kids learn versus the way you ensconce knowledge in your head. Now Tony articulates it as like yin yang or Kahneman’s [thinking] fast and slow and any of those things.
Larry:
But are you doing stuff to try to leverage the … One of the classic things is the conversational benefits of LLM, the ability and their ability to write queries because they’ve read the whole internet and have seen every SQL query ever written. Are you using those kind of capabilities in your architectures? Or …
Brad:
Yeah, absolutely. And we’ve demonstrated a couple of those capabilities. Sometimes deploying that in production requires a degree of assessment and verification that, well, takes a long time and is more complicated than one thinks. So we could name-drop a ton of people who are working in probabilistic circuits or other approaches to neuro-symbolic AI, bidirectional or … What’s the other one? Active inference solutions, where you are trying to use some of the probabilistic outcome in conjunction with deterministic functions. There’s some really interesting folks out there doing this. What we’ve seen is to have a natural language interface translate between a human description of a query and decipher representation of that query. There may be other ways to do that. Maybe text decipher is a solved problem for all I know, but that type of use case is actually really, really useful.
Brad:
But more generally, that type of use case is going to be required to solve some of the problems that our technologically enabled world has found themselves in. At the risk of going back to justice and public safety, especially for severe incidents, there can be a huge amount of video evidence submitted, digital evidence that’s submitted as video. When you’re talking about certain types of cases, that can be really extreme graphic imagery as well. And so there is a volume and scale of data now in some particular niches, maybe I’ll say, some particular use cases that may require a degree of automation outside of these existing processes. But as that translation between the human description of some goal that I have and some desired outcome, what the desired outcome actually looks like, I mean, that type of translation is, I think, what we’re going to see these systems are actually for and do the best is as that translation layer, so to speak.
Brad:
And that can be generating a Java beam that does some mapping function for you that translates from one representation to another. Maybe that’s something you do in NIEM, for example, all the time to make that localization from canonical to localized description easier to build and easier to implement into a software workflow, or maybe it is for some sort of researcher querying exploratory type of question, where you’re trying to ask questions of knowledge graphs worth of data and then the backend handler needs to know how to do all those things that you’ve asked. And that’s why you’re seeing the rise of agentic systems or bots more generally is what we used to call them, like a bot that just executes a function. You want to make the functions reliable, but your ability to call one function or another as flexible as human beings are.
Larry:
I think 2026 is shaping up as the year that where a lot of this architectural stuff is really going to … It just feels like there’s a lot going on there. But hey, Brad, I can’t believe it. We’re coming up close to time already, but before we wrap up, is there anything last, anything you want to revisit from the conversation or just make sure that you share before we wrap up?
Brad:
Well, in the future, I’d love to talk more about entity resolution and some of the great work that Paco and Senzing are doing that I think has really this profound relevance in making data more meaningful and clearer while allowing for different representations. But that determinism is more useful than probabilistic outcomes, but some appreciation of probabilistic thinking is necessary to solving the problems the world finds itself in the future. But I think, Larry, the crux of what I’d really like to emphasize in closing is that we kind of owe it to ourselves and the future of the human organism to approach these problems in a new way and to really bring an appreciation of systems thinking and systems analysis to the problems of domain-driven design and solving, making, I should say, more effective services that work together in a more ethical and effective way. And I hope to work with you and others like us who are interested in working on that.
Larry:
Yeah, that’s again, another whole episode on nerding out on systems and then systems design and domain-driven design and the intersection of them and doing it in a way that honors our humanity, which is I think is one of the things I like about our work.
Larry:
Hey, one very last thing, Brad, if folks want to stay connected, what’s the best way to find you online or connect with you?
Brad:
The best way, probably the easiest is to connect with me on LinkedIn, Bradley Bolliger. You can send me an email at bradbollinger@gmail.com, but I’m not that great at reading emails, so connect with me on LinkedIn.
Larry:
Excellent. I’ll put that in the show notes as well. Well, thank you so much, Brad. It’s always fun talking, and this was a particularly fun conversation.
Brad:
Thank you for having me. This is exciting and looking forward to seeing what comes of it.