Podcast: Play in new window | Download
Subscribe: Apple Podcasts | Spotify | Amazon Music | Android | Youtube Music | RSS

Capturing knowledge in the life sciences is a huge undertaking. The scope of the field extends from the atomic level up to planetary-scale ecosystems, and a wide variety of disciplines collaborate on the research.
Chris Mungall and his colleagues at the Berkeley Lab tackle this knowledge-management challenge with well-honed collaborative methods and AI-augmented computational tooling that streamlines the organization of these precious scientific discoveries.
We talked about:
- his biosciences and genetics work at the Berkeley Lab
- how the complexity and the volume of biological data he works with led to his use of knowledge graphs
- his early background in AI
- his contributions to the gene ontology
- the unique role of bio-curators, non-semantic-tech biologists, in the biological ontology community
- the diverse range of collaborators involved in building knowledge graphs in the life sciences
- the variety of collaborative working styles that groups of bio-creators and ontologists have created
- some key lessons learned in his long history of working on large-scale, collaborative ontologies, key among them, meeting people where they are
- some of the facilitation methods used in his work, tools like GitHub, for example
- his group’s decision early on to commit to version tracking, making change-tracking an entity in their technical infrastructure
- how he surfaces and manages the tacit assumptions that diverse collaborators bring to ontology projects
- how he’s using AI and agentic technology in his ontology practice
- how their decision to adopt versioning early on has enabled them to more easily develop benchmarks and evaluations
- some of the successes he’s had using AI in his knowledge graph work, for example, code refactoring, provenance tracking, and repairing broken links
Chris’s bio
Chris Mungall is Department Head of Biosystems Data Science at Lawrence Berkeley National Laboratory. His research interests center around the capture, computational integration, and dissemination of biological research data, and the development of methods for using this data to elucidate biological mechanisms underpinning the health of humans and of the planet. He is particularly interested in developing and applying knowledge-based AI methods, particularly Knowledge Graphs (KGs) as an approach for integrating and reasoning over multiple types of data. Dr. Mungall and his team have led the creation of key biological ontologies for the integration of resources covering gene function, anatomy, phenotypes and the environment. He is a principal investigator on major projects such as the Gene Ontology (GO) Consortium, the Monarch Initiative, the NCATS Biomedical Data Translator, and the National Microbiome Data Collaborative project.
Connect with Chris online
Video
Here’s the video version of our conversation:
Podcast intro transcript
This is the Knowledge Graph Insights podcast, episode number 37. The span of the life sciences extends from the atomic level up to planetary ecosystems. Combine this scale and complexity with the variety of collaborators who manage information about the field, and you end up with a huge knowledge-management challenge. Chris Mungall and his colleagues have developed collaborative methods and computational tooling that enable the construction of ontologies and knowledge graphs that capture this crucial scientific knowledge.
Interview transcript
Larry:
Hi everyone. Welcome to episode number 37 of the Knowledge Graph Insights podcast. I am really delighted today to welcome to the show Chris Mungall. Chris is a computational scientist working in the biosciences at the Lawrence Berkeley National Laboratory. Many people just call it the Berkeley Lab. He’s the principal investigator in a group there, has his own lab working on a bunch of interesting stuff, which we’re going to talk about today. So welcome, Chris, tell the folks a little bit more about what you’re up to these days.
Chris:
Hi, Larry. It’s great to be here. Yeah, so as you said, I’m here at Berkeley Lab. We’re located in the Bay Area. We’re just above UC Berkeley campus. We have a nice view of the San Francisco Bay looking into San Francisco, and so we’re a national lab, so we’re part of the Department of Energy National Lab system, and we have multiple different areas here in the lab looking at different aspects of science from physics, energy technologies, material science. I’m in the biosciences area, so we are really interested in how we can advance biological science in areas relevant to national scale challenges really in different areas like energy, the environment, health and bio-manufacturing.
Chris:
My own particular research is really focused on the role of genes and in particular the role of genes in complex systems. So this could be the genes that we have in our own cells, the genes in human beings, how they all work together to hopefully create a healthy human being. One part of my research also looks at the role of genes in the environment, and in particular the role of genes inside tiny old microbes that you’ll find in the ocean water and in the soil. And how these genes all work together, both to help drive these microbial systems, help them work together and how they all work together really to drive ecosystems and biogeochemical cycles.
Chris:
So I think the overall aim is really just to get a picture of these genes and how they interact in these kind of complex systems and build up models of complex systems from scales right the way from atoms through the way through to organisms and indeed all the way to earth-scale systems. So my work is all computational. I don’t have a wet lab. So one thing that we realized early on is just when you are sequencing these genomes and trying to interpret the genes, you’re generating a lot of information and you need to be able to organize that somehow. And so that’s how we arrived at working on knowledge graphs, basically to assemble all of this information together and to be able to use it in algorithms to help us interpret biological data and help us figure out the role of genes in these organisms.
Larry:
Yeah, many of the people I’ve talked to on this podcast, they come out of the semantic technology world and apply it in some place or another. It sounds like you came to this world because of the need to work with all the data you’ve got. What was your learning curve? Was it just another thing in your computational toolkit?
Chris:
Yeah, in some ways. In fact, my background is, if you go back far enough, my original background is more on the computational side and my undergrad was in AI, but this is back when AI meant good old-fashioned AI and symbolic reasoning and developing Prolog rules to reason about the world and so on. And at that time, I wasn’t so interested in that side of AI. I really wanted to push forward with some of the more nascent neural network type approaches. But in those days, we didn’t really have the computational power and I thought, “Well, maybe I really need to, I actually learned something about biological systems before trying to simulate them.” So that’s how I got involved in genomics. This was around about the time of just before the sequencing of the human genome, and I just got really interested in this area, a position came up here at Lawrence Berkeley National Laboratory, and I just got really involved in analyzing some of these genomes.
Chris:
And in doing this, I came across this project called the Gene Ontology that was developed by some of my colleagues originally in Cambridge and at Lawrence Berkeley National Laboratory. And the goal here was really as we were sequencing these genomes and we were figuring out there’s 20,000 genes in the human genome, we discovered we had no way to really categorize what the functions of these different genes were. And if you think about it, there’s multiple different ways that you can describe the function of any kind of machine, whether it’s a molecular machine inside one of your cells or your car or your iPhone or whatever. You can describe it in terms of what the intent of that machine is. You can describe it in terms of where that machine is localized and what it does, and how that machine works as part of a larger ensemble of machines to achieve some larger objective.
Chris:
So my colleagues came up with this thing called the gene ontology, and I looked at that and I said, “Hey, I’ve got this background in symbolic reasoning and good old-fashioned AI. Maybe I could play a role in helping organize all of this information and figuring out ways to connect it together as part of a larger graph.” We didn’t call them knowledge graphs at this time, but we’re essentially building knowledge graphs at the time and make use of, in those days quite early semantic web technologies. This is even before the development of all the web ontology language, but there was still this notion that we could use, we could use rules in combination with graphs to make inferences about things. And I thought, “Well, this seems like an ideal opportunity to apply some of this technology.”
Larry:
That’s interesting. It’s funny we didn’t plan this, but the episode right before you in the queue was of my friend Emeka Okoye. He’s a guy who was building knowledge graphs in the late ’90s, early 2000s, mostly the early 2000s before the term had been coined, and I think maybe even before a lot of the RDF and OWL and all that stuff was there. So you mentioned Prolog earlier, and what was your toolkit then, and how has it evolved up to the present? That’s a huge question. Yeah.
Chris:
I didn’t mean to get into my whole early days with Prolog. Yeah, I’ve definitely had some interest in applying a lot of these logic programming technologies. As you’re aware, there was different kind of branches in the development of some of these semantic technologies back in the 1990s that there was a lot of interest in deductive databases and data log systems. But for whatever reason, the Semantic Web decided to adopt more description, logic type technologies. And at the end of the day, a lot of these things, they’re quite similar in the overall goals. You’re trying to state your knowledge about some of the rules of these kinds of systems.
Chris:
And so for a biological system, this could be describing rules about how different parts of the cell pieced together and how things are assembled. So the cell wall forms the outer part of the cell, the plasma membrane surrounds the inner contents of that cell. So you can start making use of some of these systems, whether it’s Prolog type rule systems or description logics to encapsulate some of that knowledge about cell biology. And then make use of reasoners or problem or constrained satisfaction engines to essentially make inferences about those systems to see if some of the statements that you’re making, if they’re logically consistent or if one kind of statement can follow on from another.
Larry:
Yeah, we were talking a little bit before we went on the air about how most of my conversations have been with people in the enterprise and the business in other worlds, a little bit with pharma and things like that. But the way you just described your need to capture knowledge about cell biology is exactly analogous to the things that here in that world. So it gets us to the practice of building these things. I mean, the title of this podcast is Knowledge Graph Insights, so I’d really like to focus on that.
Larry:
Getting to that point though, one of the things I really want to talk to you about is, so there’s all this innate complexity that you’ve already talked about, just the insane number of genes, and that’s just the human genome, and then there’s all these other genomes as well, and that’s in the context of this giant complex field. Plus there’s all kinds of different, you mentioned adjacent professions and things like that, adjacent disciplines, there’s all these collaborators. Just the simplest thing, how were these ontologies built? I’m imagining there was a little bit of collaboration somewhere along the way.
Chris:
Absolutely. Yeah, no, there’ve always been collaborations. And one thing that I think distinguishes the biological ontology community from maybe some of the other communities building ontologies is this has always really been driven by the biologists and been driven by bio-curators. So this is a profession that we have within biological information systems just because the sheer scale and complexity of the knowledge that we need to capture is so wide-ranging. So there’s actually a number of different bio-curators employed globally within the US, UK, Japan, other places, whose job is really to sift through the literature and try and capture these nuggets of information, nuggets of knowledge about genes or about diseases or about chemical entities and so on. And capture them in structured forms, whether that’s in a relational database or whether it’s in an ontology. Our community has really been very driven by these biologists.
Chris:
Early on we realized we needed to bring together some of the expertise in knowledge representation and computational biology with some of this more kind of biologically driven knowledge representation. But these ontologies, a lot of them initially started out with a handful of curators really working together using version control to build up an ontology. But pretty soon we start to realize that none of these ontologies are entirely independent amongst themselves within the gene ontology. We have a need to be able to talk about cell types or to be able to talk about different kinds of chemicals that are produced or consumed by the cell. And these all live in their own separate independent ontologies, but we need to figure out ways to connect between them and form bridges between them. So we have this challenge of being able to collaborate within our own individual domains, but also speak the same language as our collaborators.
Chris:
And at the end of the day, this just involves a lot of sitting down and working with people and trying to understand their worldview and trying to align that worldview because the way a chemist might think about how to classify chemicals or how to interlink them together might be different from a biochemist or a medicinal chemist. So these things just, they don’t all link together automatically just based on us using the same terms. And there’s a nice paper by Janna Hastings and Fabian Neuhaus, and really the title of the paper is something like Ontology Development is Consensus Creation, and it’s this idea that we’re not really just sitting in isolation working on our own individual ontologies. But really the job of an ontologist is to bring together these experts in the domain and experts in knowledge representation, and try and forge some kind of consensus community view of how to represent that knowledge.
Larry:
In my mind, there’s a little bit of a chicken and egg here thing right now. You mentioned the bio-creators earlier are these people who have been curating and scanning the literature to capture all the concepts you’re working with. And then we introduced the role of the ontologist. Were there ontologists who were kind of guiding those folks? Because a classic thing in ontology design is the ontologist kind of guiding or coaxing or whatever they have to do to get that knowledge from the subject matter experts. Can you talk a little bit about that relationship?
Chris:
Yeah, I think, and again, I’d say a lot of our ontologists were originally bio-curators. And for us, the role is often somewhat fluid, and people often wear both hats at the same time. And ontologists are really a collaboration between these different groups. Different groups have different models for working with these kinds of expert communities. Sometimes this is more ad hoc. We’ll realize that within a particular ontology, maybe there’s an area of interest that we have gaps in. So we got and we identify experts, or maybe we just meet them at the conference and say, “Hey, we working on the gene ontology that we have this portion of the ontology we’d like to expand out and flesh out a little bit better. Maybe you and some of your colleagues could help.” Some groups have started developing more structured ways of doing this, making use of various kinds of methodologies to help in doing this.
Chris:
But really it varies across different groups. So maybe to give one example in one of the knowledge graphs that I work is about representation of diseases and how diseases are classified, how those diseases relate to the underlying genetic causes of those diseases, how they relate to individual symptoms of those diseases. And we were working on one part of the ontology to do with epilepsy, and we realized we needed to modernize the classification we had there. So we organized a series of workshops. We brought together experts, not only clinical experts in epilepsy, but also geneticists and even patient advocacy groups as well. Because we think it’s important for everyone to have a voice in how diseases that they either work on or that they even suffer from are classified in ontologies and how they’re linked together with other resources.
Larry:
That’s really interesting. And you can kind of picture how the exact structure of any one cohort like that might vary. You mentioned, I think when we were talking to prepare for this, you mentioned the Delphi Method that came out of the RAND Corporation, and you just alluded to others. Are there sort of, I don’t know, I don’t like the word best practices so much, but are there better practices and better methods by which these highly collaborative projects have come to fruition?
Chris:
There’s absolutely many lessons learned from undertaking this process multiple times in different ontologies. I think one of the key lessons is really just to meet people where they’re at. You can’t just sit a domain expert down in front of an ontology development environment like Protégé. We love Protégé for our own use in developing ontologies, but it’s quite a technical tool. And so really you have to figure ways to communicate the structure of the ontology that doesn’t simplify it too much, but allows people to work where they’re at. And some people are making use of things like the Delphi method in combination with various kinds of workflows and specialized user interfaces that allow people to vote on different ways of representing things. But often at the end of the day, it comes down to just Google Sheets and spreadsheets and sharing information with people in simple tabular forms that they’re familiar with. Unfortunately, that’s where we’re at a lot of the time with our tooling.
Larry:
I have navigated many different professions and talked to a lot of people about this, and the answer is always spreadsheet, whatever your data wrangling problem is. Yeah, no, I wish I’d invested in any, but yeah, no, but at the same time, that kind of gets at. So there’s this huge diversity of different ways that people work. I was talking to a friend the other day who works with just oceanographers and marine biologists, and he was talking about all the different ways, different baselines they have and different measurement technologies and all that stuff. So just normalizing the data and labeling things in just that one little tiny field and what you just said. But I think for one of these domain experts who are practitioners out there, they’re probably using something like a spreadsheet that works for their needs. What are the high level lessons about how you understand these things coming from a variety of sources in a way that lends themselves to a discipline-wide ontology?
Chris:
Yeah, I know that’s a kind of big overarching question, but I think a lot of it comes down to having the right terminology to express your ideas as well. I mean, I think ontology 101 is learning that the actual labels that you use for things shouldn’t really matter. It should all be about the concept. You can apply whatever terms that you like, but then the rest of actually applying to ontologies is actually working with communities and using their terminology with all its individual nuances and being able to communicate across these different ways of talking about the same thing.
Larry:
Is there a Rosetta Stone? In the case of a RDF triple store based knowledge graph? You have a W3C standard to anchor to. Are there other standards or things that help you align folks on how to do stuff?
Chris:
Yeah, I guess we’re not aligning at the level of those standards themselves. I mean, at the base of most of our ontologies is OWL and the web ontology language that is driving things. Even if we do abstract away from that, we’ve developed tooling that allows people to use things like spreadsheets and then have it essentially compiled down into OWL, ways of integrating these workflows with using GitHub and things like that as well. So I think it really depends on the immediate community that you’re trying to reach and the appropriate level of standards to expose them to. I guess.
Larry:
Hey, and I want to go back to the people stuff a little bit because that’s the thing that’s really intriguing to me about this is its such a diverse group of folks, and I assume it’s global as well, that there’s scientists all around the world. Yeah, I guess, and I know that these are just well-established disciplines that have big 10, 20, 30,000 person gatherings and stuff like that. Is most of the collaboration happening just on Zoom like we are right now? And in particular the ontological collaboration? Are there, I don’t know, I’m thinking of workshop-ey facilitation kind of methods and things like that. Are there practices around that that have stood out for you?
Chris:
Absolutely. Yeah. As you might expect, a lot of this is in-person meetings, but especially during the pandemic when that was harder, we switched to a lot more Zoom meetings. It’s also a lot more expensive and hard to coordinate to bring people together in the same location. So a lot of it comes down to the actual interaction with people is incredibly important, face-to-face interaction like we’re having right now. But a lot of our work really has been transformed a lot by adoption of really social coding mechanisms. So things like GitHub. So we switched to using GitHub for a lot of ontologies I think over about 10 years now. And some people were a little bit reticent at first. GitHub is, it was perceived to be designed for software developers. It had a very quite techy way of interacting with it at first, Git is not the easiest system to learn, but we felt the advantages of the social aspects of that mechanism were so powerful that we really wanted to make a push to get everyone on board and start using that platform.
Chris:
And that’s turned out to be really transformative for us because before then, a lot of decisions were made by people perhaps emailing each other or one-on-one conversations, and we weren’t tracking those and we weren’t tracking the reasons for our decisions. And what would happen is five years later, we’d come back to the same area of the ontology and try and refactor it and then realize that, wait a minute, didn’t we previously work on this five years ago, aren’t we reverting back to what we did before. So having this public, transparent, open mechanism where we can track all of our discussions and tie those discussions to the individual changes that we’re making in the ontology has turned out to be really incredibly helpful.
Chris:
And we’re really seeing that now when you got this ten-year window, when you can look back and see the history for all of the decisions that you’ve made. And it also kind of helps the way that you can have these kind of interlinkages between some of these repositories as well. And I should mention also that we made the decision right from the early days to adopt version tracking, version control methodologies from software engineering. We felt it was really important to treat change in our ontologies as essentially this first-class entity that we can easily track. And also it was lightweight as well. We didn’t have to build our own infrastructure. We could just take advantage of all of this existing infrastructure.
Larry:
When you think about all the change over the last 10 years, just the mass adoption of Git and that in the sort of built-in, not requirement, but it’s sort of self-documenting at least when decisions were made. And there’s probably somebody did some kind of little commit note about it. Yeah. Now that’s fascinating. What about the, just I’m thinking now of with these broad collaborations, I know within any one profession or discipline, it’s always you’re generally speaking the same language, but I know there’s also a lot of cross-disciplinary stuff. In terms of the people organizing part of it, is that a challenge at all? You mentioned biochemists and marine biologists or something looking at the same terrain a little bit differently. Do you have to navigate that stuff very often?
Chris:
Absolutely. Yeah. There’s both the reaching consensus about our worldview and how we’re modeling the domain. But I think even more importantly is we bring to this a lot of tacit assumptions about why we’re building an ontology in the first place. So for molecular biology, we’re often building these ontologies for purposes of analyzing data. One of the main use cases for the gene ontology, the cell ontology, is now that we have all of these genes annotated using these ontologies, we can take a lot of complex experimental data where you see many genes being over expressed or under expressed and changing response to different conditions. And it’s impossible for a human being to look at that list of genes and figuring out what’s going on.
Chris:
So we make use of these ontologies to help us interpret and analyze that and say, “Okay, yeah, I can see that in response to this condition, this particular biochemical pathway or biological pathway has become activated, or that this cell type plays a really important role in this disease. But then we bring some of our tools and ideas to maybe a neighboring community and they don’t have the same use cases. They want to use an ontology to structure the metadata that they’re capturing up-front about their experimental systems or their assays. And just those different perspectives really bring to bear just a lot of different approaches and how you might go about building those ontologies and what you might prioritize and who you might talk to and going forward with this.
Larry:
Hey, you’re blowing my mind a little bit with it because I have a number of friends who are scientists and some of them spend literally years just building the equipment they need to measure the thing. But what you’re saying about that, but just those two use cases you just mentioned, like analysis of this well-documented genome versus structuring metadata about how you do stuff, each of which is valuable and probably equally useful. That’s amazing. But hey, Chris, I can’t believe it. We’re coming up close to time already. But before we wrap up, is there anything last, anything you’d want to revisit from the conversation or just make sure we share before we wrap up?
Chris:
Yeah, no, I think we covered a lot of the areas that I’m interested in the moment, but I guess I can’t really leave things without talking about some of what we’re doing at the moment with making use of AI and particular agentic technology like everyone else when ChatGPT came out in 2023. We started using it to build toy ontologies, and it was good for that, building toy ontologies, but just using them in isolation, wasn’t really particularly useful. And there was a bit of a misalignment between the kind of things that an LLM could do and the more complex socio-technological things that we really need help with with ontology at buildings. So right now we’re having a lot of fun playing with agents, playing with multi-agent systems. And again, we’re using some of the same methodologies that have worked for us before where we have a look and see what our colleagues and software development are doing.
Chris:
Before it was like using version control, using Docker to containerize CI/CD. Now we’re seeing, hey, a lot of software developers are vibe coding using tools like Claude Code. Maybe we could have some of that. So we’ve started building up some patterns and some best practices about how you can take some of these tools and use them not just for building code, but for building ontologies as well. And because we have this fantastic version history of everything going back in GitHub, it actually makes a really nice set of benchmarks and evaluations that we can use to do this. So we’re currently just kind of exploring all of the different ways you could have some kind of multi-agent system with different kind of specializations in the domain science in logical reasoning and how you can bring those together and use them to help us build ontologies.
Larry:
Oh, man, that’s a whole other conversation I want to have now because that’s… and one of the things I didn’t have on our list for today was you’ve done so much in terms of tools development and the infrastructure for doing all this stuff. And it sounds like this is just exploding with the arrival of ChatGPT and LLMs and all that stuff. So that’s awesome. A quick follow up on that, of all those experiments you’ve done, is there anything that’s particularly promising or useful in your work?
Chris:
Well, as you might expect, the AI does fairly well with fairly simple kinds of requests where it’s like, “I want to add a new term into an ontology.” But where it’s actually kind of surprising us, it’s doing quite well in a lot of harder tasks as well. And where I think is really going to be helpful is we just have a lot of really tedious refactoring. It’s just like software development. You build a software library over 10 years or so, it’s vital that you need to maintain that to keep it running. But there’s all kinds of areas where you want to go back and refactor parts of it and make changes, and it’s turning out to be really useful for those kinds of larger tasks as well.
Chris:
And so one thing we did apply it for, which was a fairly simple task, but was something no one really wanted to do was many of our ontologies include provenance for information in the ontology, making use of URLs from external sites. So this might include for infectious diseases. We made use of a lot of URLs from the CDC. More recently, those URLs have started to break a bit more frequently than they did in the past. So it turns out that the agent is quite good at going back, going through analyzing, looking at where those URLs are broken, finding good potential replacements that they can then suggest to the curator where we can then go back and fix some of those URLs. So we’re using it for some of these tedious tests as well as some of the bigger logical reasoning type. Harder tests too.
Larry:
That’s great. It’s just great how these hybrid systems are the complimentary natures of these technologies, permit all of this. You can do a lot more it sounds like now, so that’s awesome.
Chris:
Absolutely. Yeah. No, and the idea is really to provide assistance to the experts who are building this and help clear their plate to have some of the more tedious mechanical tasks and allow them to focus on the real high level aspects of this.
Larry:
Nice. Hey, one very last thing, Chris. If folks want to follow you or connect online, what’s the best place to find you?
Chris:
My name is pretty Google-able, so you can find my institutional webpage, but I think probably the best is just my LinkedIn page. I guess we can drop a link for others to see.
Larry:
Yeah, I’ll put that in the show notes for sure.
Chris:
Okay.
Larry:
Well thank you so much, Chris. I really enjoyed the conversation.
Chris:
All right, well thank you, Larry. It’s been great.