Podcast: Play in new window | Download
Subscribe: Apple Podcasts | Spotify | Amazon Music | Android | Youtube Music | RSS

In recent years, data products have emerged as a solution to the enterprise problem of siloed data and knowledge.
Andrea Gioia helps his clients build composable, reusable data products so they can capitalize on the value in their data assets.
Built around collaboratively developed ontologies, these data products evolve into something that might also be called a knowledge product.
We talked about:
- his work as CTO at Quantyca, a data and metadata management consultancy
- his description of data products and their lifecycle
- how the lack of reusability in most data products inspired his current approach to modular, composable data products – and brought him into the world of ontology
- how focusing on specific data assets facilitates the creation of reusable data products
- his take on the role of data as a valuable enterprise asset
- how he accounts for technical metadata and conceptual metadata in his modeling work
- his preference for a federated model in the development of enterprise ontologies
- the evolution of his data architecture thinking from a central-governance model to a federated model
- the importance of including the right variety business stakeholders in the design of the ontology for a knowledge product
- his observation that semantic model is mostly about people, and working with them to come to agreements about how they each see their domain
Andrea’s bio
Andrea Gioia is a Partner and CTO at Quantyca, a consulting company specializing in data management. He is also a co-founder of blindata.io, a SaaS platform focused on data governance and compliance. With over two decades of experience in the field, Andrea has led cross-functional teams in the successful execution of complex data projects across diverse market sectors, ranging from banking and utilities to retail and industry. In his current role as CTO at Quantyca, Andrea primarily focuses on advisory, helping clients define and execute their data strategy with a strong emphasis on organizational and change management issues.
Actively involved in the data community, Andrea is a regular speaker, writer, and author of ‘Managing Data as a Product’. Currently, he is the main organizer of the Data Engineering Italian Meetup and leads the Open Data Mesh Initiative. Within this initiative, Andrea has published the data product descriptor open specification and is guiding the development of the open-source ODM Platform to support the automation of the data product lifecycle.
Andrea is an active member of DAMA and, since 2023, has been part of the scientific committee of the DAMA Italian Chapter.
Connect with Andrea online
Video
Here’s the video version of our conversation:
Podcast intro transcript
This is the Knowledge Graph Insights podcast, episode number 30. In the world of enterprise architectures, data products are emerging as a solution to the problem of siloed data and knowledge. As a data and metadata management consultant, Andrea Gioia helps his clients realize the value in their data assets by assembling them into composable, reusable data products. Built around collaboratively developed ontologies, these data products evolve into something that might also be called a knowledge product.
Interview transcript
Larry:
Hi, everyone. Welcome to episode number 30 of the Knowledge Graph Insights podcast. I’m really happy today to welcome to the show Andrea Gioia. Andrea’s, he does a lot of stuff. He’s a busy guy. He’s a partner and the chief technical officer at Quantyca, a consulting firm that works on data and metadata management. He’s the founder of Blindata, a SaaS product that goes with his consultancy. I let him talk a little bit more about that. He’s the author of the book Managing Data as a Product, and he’s also, he comes out of the data heritage but he’s now one of these knowledge people like us. So welcome, Andrea. Tell the folks a little bit more about what you’re up to these days.
Andrea:
Thank you. Thank you very much, Larry, for having me. It’s a pleasure. Yes, as a CTO in Quantyca, I’m in charge of all our advisory services. So I’m helping a customer in figure out how to manage their data properly, especially to leverage the potential of artificial intelligence. So basically I see all sort of problem in data management. Each client, it’s different, but each client have a lot of problem of data that is very fragmented or too complex to manage. And so it’s a very complex problem to feed this data to the AI model and extract the potential that the modern AI and all the breakthroughs that we are seeing in this day made available. So I’m really focused at this moment to help customers, especially in find a way to manage their knowledge, the knowledge that is characteristic of the company, that is a differentiator of the companies, the knowledge that is not known at the large language mode, what make the company different and can be leveraged to implement domain-specific, company-specific use case based on AI and leveraging the data collected.
Larry:
Yeah. As you mentioned that, we were just chatting a bit before we went on about the scope of the conversation. And I totally forgot to mention AI, which is of course is like the main driver for half of this stuff we’re doing nowadays. But a couple of things you mentioned there. I want to go back to, one, you mentioned what a complex problem space this is and the challenges of data management and every organization has its own issues there. One of the ways that folks like you have helped people cope with this is the notion of a data product. And I know that’s a newish concept and maybe new to some of the listeners to this. Can you talk a little bit about your conception of what a data product is and how you put one together?
Andrea:
Yeah, absolutely. The concept is new but the rationale behind it, it’s not new. Humans, when a problem is too much complex, the only way that humans have found to solve a very complex problem is to split in a different part, in smaller part, and try to take all the complexity within each single part. And the idea of that a product come from this strategy, done the lead, the team part. So the idea is not managing the data in a unique central platform in which all the data of the company is collected but split in a modular architecture. So the platform is still there. You have the data layer, the data warehouse, whatever. It’s the architecture that you prefer, but it’s not anymore a monolithic solution in which you store all the data that you have in your company, but it’s built as a composition of independent modules. Each module focuses on one or more, but usually one specific data asset, and there is a team that is in charge of manage the life cycle of that data product that manages specific data asset.
Andrea:
Of course the composition of all the data asset create the platform and the platform can be used to support the different use cases, but basically you can work on each single module without caring too much about the other module because each module is isolated with a specific interface. So if you do not modify the interface, you can modify the technology and implementation inside. And if you want to understand how the different modules connect, you can ignore the implementation and just concentrate on the relation between the different interfaces.
Andrea:
So to make it very, very simple, we can think at the data product like a sort of microservice that is a software application, is actually a software application, that does not expose functionality, transactional functionality, to acquire data and drive the transaction but expose the data. It’s a software application that expose the data in order to make the data it manages as much usable as possible for its customer base, for its users. So this is a data product. And of course because it is a product, it is managed with a product mindset. So it’s not a project. It’s not something that the team develop and then forget about it. But there is a dedicated team that implement the first version and then evolve the software application that support that specific data asset through all its lifecycle till the retirement when the data asset is not anymore relevant for the company. That’s pretty much what is a data product for me.
Andrea:
So basically I call this kind of data product a pure data product to even more underline the fact that it’s a software application that expose data because I also have a lot of time the question, a report, a dashboard is a data product and they say, yes, it’s a data product if it is managed as a product with a product mindset. But my book, my research is more focused on the pure data product, so that specific kind of data products that do not expose visualization or insight or action but expose just pure data to make it reusable and composable over time to support multiple use cases.
Larry:
That’s right. We didn’t talk about this before we went on the air, but the episode right before this one is with Dave McComb, and I know I’ve heard you talk before about you appreciate his approach and his data-centricity. And everything you just said, I’m like, “Oh yeah, he’s read Dave’s books.” Was that the major influence, or what are the influences?
Andrea:
Absolutely. It was for me an epiphany because at that time when I read McComb’s books, I was looking for… I had a problem because we had started since couple of years to help our customer and created this kind of modular architecture. So that architecture that is built as a composition of different data product, even managed with a distributed operating model. So all the data product are managed by different business domain in an autonomous way. And we have seen that this improve a lot the quality of the data and reduce the maintenance cost because of course it’s a better way to manage the complexity of the whole solution.
Andrea:
But we also found out that the data product portfolio is not reused a lot. So we created good data product, but this data product are usually not composed with data product belonging to different domains or reused for multiple use cases. And so a lot of the value of managing data as a product was left on the table. And so I was looking for something to resolve this problem to improve the composability and the reuse of the data product because at the end of the day, you invest in data product, not just because you want to have more quality data or to reduce the maintenance costs are fantastic things of course, but also because you want to have a library that you can reuse over time and reduce the integration. The more use case you implement, the more data product you have in your library. And the biggest is the probability that the new use case can be implemented without doing any more integration, but just composing already existing data problem.
Andrea:
So I was looking for a solution for this problem when I read the book of McComb and I say, “Okay, probably an ontology in the center of the company that define the common language that can be used to compose not only from a technical standpoint, but also from the semantics standpoint, the different data product is exactly what I am looking for.” And so I start to move on on the ontology field on the knowledge field and study and find out how to create this ontology and how to connect that product through the ontology. But it was a real inspiration because when you have a problem and you read something that say this is the solution we’re looking for, probably I have to go this way and was an epiphany for me.
Larry:
That’s right. And you had to read both of his books because the second book is the solution. The, what is it, data center?
Andrea:
Yeah, I read the red one first and I feel the pain and said, “This is my problem. This is my problem. This is the problem.” The solution then, the green book, “Okay, that’s a solution.” Not sure that it’s viable, but I want to invest in study this thing because it seems it makes sense to me. It resonates a lot with me with my problem. So I arrived at the ontology world not because of large language model. The ChatGPT was not released yet at the time. It’s not for me an interest to reduce hallucination in large language model. My interest that take me to ontology is how to compose product that are developed by different domains.
Larry:
And that notion of a lot of stuff you just said reminds me of, well, first of all, that notion you just said of domain-focused stewardship of the data and the data products. But also you talked about just generically some of the highest level concerns that enterprise architects have now this notion of composable architecture and these modular designs that permit that. And then with the intent of that being so that you can reuse stuff, because that’s, I think, one of the things that drives Dave nuts and everybody like you, like, “Oh really, you’re going to reinvent the wheel with this new application?”
Larry:
And so I guess you’ve just outlined a lot of the generic benefits of a data product, the reusability, the composability, and something else you just said in there too. You mentioned use cases. And I don’t know if, I can’t remember if this… Yeah, this one’s out now. I had Jacobus Geluk on a while back and he’s talking about he’s got this new idea of use case trees. But I think that’s one of many models by which you can account for, and you’ve already mentioned a couple of examples of how you account for various use cases. So can you talk a little bit about that, the uses that these data products are serving?
Andrea:
Yeah. Basically, each data product is a software application is concentrated on a specific data asset. This data asset can be a portion of the data related to the customer, a portion of the data related to the contract, with the product, whatever is important in your company, the business object that are important in your company that generate data through different processes in your company. And the idea is to expose this data in order to make it super simple to be used. So it’s not modeled, the data, to just answer to one specific use case and make it more complex to be a reused in another use case, but just exposing the data in a general way, in a way that makes sense for the company that describes well that specific data asset.
Andrea:
And so I can use it for one use cases, but in the future I can reuse it for multiple use cases. Because in a company for example, I will have in the future a lot of use cases that need to get the data related to the customer, for example. So if I have a good data product that expose the data of the customer, of course not all the data of the customer because the customer is a super big entity, but maybe the data related, the sales attribute of the customers, or where to send invoices or whatever. All the use cases that needed this data have not to rebuild the integration, have not to do the work to create that integration, to take this data from source system but can rely on one specific data product that every team in charge that is there to expose this data in order to make it more usable as possible for different use cases.
Andrea:
This is very important because if data products do not create this level of reusability for multiple use cases, basically a data product in a sense. It’s another way to create a data series because every time I have a new use case, I needed to create new data products. So basically there is a mapping one to one between use cases and data products. And so basically a data product is a silos in a way. It’s not a silos if I reuse it in a multiple use case, and I do not create a copy of the same data asset because I know that when the data, it’s related to the customer and it’s data related to the sales property of the customer, the data product that is in charge to manage this kind of data is the product, A, managed by team X? And they have to do that.
Andrea:
So they have to integrate new data of this kind. They have to model it. They have to expose it. They have to take accountability about the quality of this data and all the stuff related to make it a product, not to turn it into a product that consumer like to use.
Larry:
Yeah. As you’re talking about that, you’re reminding me you’ve used the term asset to refer to data products and data. And I think is that how you turn data into, from this pile, I’m picturing an oil refinery now because the data is the new oil and then you refine it into this product and that’s the valuable thing. A big pile of oil is just something to be processed, but a can of gasoline or… So that’s how the asset and that… One of the things I really like in the book is your take on that data information knowledge going up the pyramid. That’s how you turn data into an asset. Is that right? Is that-
Andrea:
Yes, data by itself is an asset because the definition of an asset is a resource that is owned by the company and that have a potential value in the future. But data is also very particular kind of asset because data has value only in use. So data by itself has zero value, acquires value when it is used. In order to make data usable, you need to describe this data. So you need to manage the metadata. And the metadata is the second level of the information, the information architecture pyramid. Because when you add to the data the metadata, you are creating information that is understandable. And then you have also to move to the knowledge level because, okay, if you have information, to understand the information in the context of your company, you needed to have a conceptual model that describes how your company works.
Andrea:
So all these three layer are important. If you have only data, data have no meaning by itself. No, it’s not understandable. And so it’s not usable. If I told you 15, 15, it’s a number. It’s a data, of course, but you cannot understand what this data mean and now you can use it. So I have to put some metadata. 15 is the sensor temperature today in LEM. So you start to understand. And then I can move to the knowledge level and say, okay, but 15, for our company it’s important because we sell ice cream. So maybe it’s a good moment to do a promotion on ice cream because it’s a good temperature to start doing this promotion and not maybe hot soup because it’s too hot. So all these three layer basically create more context around of data, and the more context you create, the more the data became usable. And the more data is usable, the more data as an asset became valuable.
Larry:
Yeah, that’s one of the best explanations of that, the ascension up the pyramid. As you’re talking about that, you mentioned the importance of metadata. And you can get metadata from relational database table or other ways of showing that, but because this is the Knowledge Graph Insights podcast so we got to talk about ontology and knowledge graphs a little bit anyway. So can you talk about the role of an ontological… Well, I guess one of the benefits of knowledge graphs and ontologies is that you can co-mingle metadata and data and do more with that. Can you talk a little bit about your take on that?
Andrea:
Yeah. Basically, I divide the technical metadata and conceptual metadata. Now, conceptual metadata is all the stuff that you put in the ontology. Technical metadata, it’s the description of the type of the field, the date when the field has been recorded, and all the kind of stuff that are more related to the logical model, let me say. Maybe not physical, but logical, not conceptual.
Andrea:
And in my area, when you manage data as a product, the team that manages the data product is accountable for the data and the technical metadata. So it’s not possible to publish a new data product if it is not annotated with all the technical metadata that make it possible to use the data contained into the data product from a technical standpoint. Instead, the conceptual model, the ontology is a way to represent the conceptual model is something that is not a responsibility of the data product team, but is managed in a central way, federated way, I like to say, because it’s the common language that must be shared by all the data products. So when I create a data product, I expose the data asset. The data asset maybe can be a table, for example. And I have to specify as a team that developed the data problem the scheme of the table, the column names, the type of the columns, and then I have to specify a link, what they call a semantic link, and say, “Okay, this table is associated to this concept in the central ontology.”
Andrea:
So if you want to get the meaning, the common meaning that the data inside this table for the company, I specify that this table is associated to this concept in the central ontology that is common to all the company. And so you know that each row in my table, it’s an instance of that concept. The concept can be for example, the customer. So each row in the table is an instance of a different customer. As the product owner, I have the responsibility to create this link. But the definition of the concept of what is a customer for our company, at least at the minimum common level that makes sense, is something that must be is a centralized object that have a life cycle that is independent from the life cycle of each data product.
Andrea:
Of course, cannot be implemented in a centralized way because it’s too much complex and it’s very difficult to manage it at a scale if there is only one team that is responsible to create this model. So I prefer to apply a federated model in the development of this enterprise ontology. But basically the enterprise ontology is the common language that we can use to map our data and to make this data interoperable not only at the technical level but also the semantic level. So if I have to join, for example, two data product that expose data related to the customer, one belonging to the marketing domain and one belonging to the sales domain, thanks to this semantic link into a common concept in the enterprise ontology, the consumer can know that the field address in one data set and the field address in the other data set are exactly the same or maybe are different, have two different meaning because one is the shipment address and the other one is the invoicing address.
Andrea:
So I can map semantically, not just make a technical join, but I know if the different field have or have not the same meaning as some can put together or not.
Larry:
Interesting. Now, you’re making me wonder. I know that this manifests differently, like the management in the organization of this. In every organization it’s handled a little differently. Are you starting to see patterns emerge of where that federation… Because the domain-specific data stewardship that makes sense and the applications and the analytic stuff over there, but where does that federation usually happen? Are there typical departments or roles where that responsibility typically lies?
Andrea:
No. When I start to propose to customer to use the ontology to create this semantic interoperability between data product in order to promote the composability of the data product, I started in a very naive way. And I say, “Okay, you already have a team of data governance, a central team of data governance. You already have a team of data stewardship, standard stewards. Why not using this team to create this central ontology in a central way?” So they are responsible to create the ontology.
Andrea:
But I see that this model basically didn’t work because first of all, it’s an organization is very complex and it’s very difficult to have in one team the knowledge of all the concept weaving the scope of the organization and tend to became more complex and the more concept you add, because every time you add a new concept, you have a lot of relationship. And the definition of complexity is not about the number of concept, but how many correlation you have between these concepts. So basically I see resurface at the conceptual level, at the knowledge level, the same problem we had in the past, in the data level, not monolithic solution managed by a central team that became a bottleneck and make not scalable the deliver of that partner.
Andrea:
So I say why we cannot apply the same ideas that we are using right now at the data level. So modularize, manage the data as a product that create a modularized solution with different team that can work in parallel. Why we cannot do the same thing at the knowledge level because we had the same problem? We have a monolithic solution managed by one team that is not able to scale as fast as the requirement that come from the organization. And I say, “Let’s try. Let’s experiment this concept of a federated operating model in which we define a different knowledge that are orthogonal to the business domain.” So the business domain can be the marketing, sales, logistics, whatever. And knowledge domain can be the customer, the product, the order, so object or key processes that are basically orthogonal, traversal to different business domain or to different value streams.
Andrea:
And we can have a dedicated team for each product in these different knowledge domains. So a team that for example, develop and maintain and avoid over time the ontology that describe what is a customer for our company or what is a product for our company or what is a contract for our company. And they are in charge and work in a federated way because then of course all these sub-ontology must be interoperable, must be linked together because the customer is one that buy a product and buy a product to making a contract. So there must be an interaction between these different sub-ontology, but each team weaving its sub-ontology can work in autonomous way. So applying even here, this principle of splitting different parts and object in these two contexts to be managed by one single team.
Larry:
And you use the phrase knowledge as a product in your book, and that’s what you’re talking about there is you get to that same… That’s really, that’s I think… And the way you just described that, the federation model, the team that manages the federation, that doesn’t have to be a very big team. Are they just setting standards and governance frameworks and things like that, or…
Andrea:
Yeah, I think that basically a team cannot be too big now because otherwise it’s not a team. What make a team a team is the fact that there is trust and alignment on purpose and vision within the team. So the team cannot be bigger than nine or between nine and 15 person. More than 15 person is not a team. It’s a group but it’s not really team because there is not the cohesion that make it really team. So it must be a small number of person, federated, maybe a person that are not full-time dedicated to the modeling team. They work within different domains.
Andrea:
So for example, if I had to define a model team that is in charge to create the ontology that describes what is a product for our company, probably is a good idea to have in that team a person that come from marketing, a person that come from sales, a person that come from operation that are not fully dedicated modeling but they periodically meet and try to find an agreement on how to define the customer. And once it is defined and review this model if it is needed over time to support new use case or because there is new needs or because there is new processes developed within the company, but it’s about to find agreement between business people that use that particular concept within their domain, in day-by-day work within their domain.
Andrea:
And then you can have also a more specific team that define basically the rule of the game. So how we define the ontology, what kind of predicate we use. There is an app at ontology that we use for all ontology, but these are just the policies. What are the rule of the game, how we model conceptually object. But the modeling itself is made by a dedicated team on dedicated knowledge product with a specific scope.
Andrea:
And the fact that is a product is another important thing because I have SOC in a lot of situation that ontologies where the enterprise ontology were defined in the top-down. So there is a team that defined ontology, but then this ontology maybe is very good but it’s not used on the field. It’s not used into the company, or just by some specific use cases but it’s not widely used within the company.
Andrea:
The fact that the knowledge product is a product, it’s designed for their user. So I have no sense to design an ontology or model a concept if there are not use cases, if there are not user that want to use that, use it. Adoption is very important, designing an ontology. So the ontology is not designed to create a perfect model of something. We are not modeling for the modeling’s sake, but we are modeling because we have interaction problem between different area of the company that speak a different language and need to find a common language to interact and to put their data together.
Larry:
I was just saying we’re coming up on time and you just introduced a whole new conversation we can have around the people part of this and the team, organization, and all that. But yeah, I’d love to return to that, but right now we’re running out of time. But I want to make sure that before we wrap, I give you a chance. Is there anything last, anything you’d like to revisit from the conversation or that you want to just make sure we share before we close?
Andrea:
No, I think that this is probably a topic for another episode, but this is for me the most important problem. The most important problem is that modeling is about people, not about the model. It’s about people. Because you have to… In my experience, people that work in different domains, maybe on the same concept, working in the same company with the same culture, see the same object in a very different way. Because a marketing person see the customer in a way that is very different from a salesperson.
Andrea:
So modeling is first and foremost about find agreement, compromise. And because if you are not able to find an agreement and you create a model that contains very little properties, have no sense because you have a language with a small number of words. So you cannot create interoperability. And if you create a model that tied to fit all the possible differences without finding agreement, just adding all the attributes or the idea that come from the different domain, you create a monster that is too much complex and no sense.
Andrea:
So you have to find the right agreement, find agreement between the different model of the world that the people have. It’s very difficult. It’s very difficult. Sometime when I also ask the same department, “Tell me what is the customer, but not loud. Write on a paper what is the customer for you?” I come out with 10 different definitions. So my work there is more with person that with the different technicalities on how to build an ontology RDF or what. It’s not that problem. It’s to make people not fight and find an agreement that is usable, but have value for the company because it’s used to put the data then together.
Larry:
Yeah, that’s right. Because I’ve had that same experience. So once you do all the hard people stuff, the model is just writes itself. But there’s a lot of work to get to that point. Hey, one very last thing, Andrea. If folks want to connect or follow you online, what’s the best place to find you?
Andrea:
For sure, the best place is LinkedIn. I’m really active on LinkedIn, so you can follow me. Usually I post with an hashtag that is TheDataJoy. Gioia is joy in Italian. So you can follow my hashtag. Or if you have any question you can answer, write me a direct message on LinkedIn and I usually answer quite quickly.
Larry:
Cool. I’ll put that in the show notes as well. Well, thank you so much, Andrea. I really enjoyed the conversation.
Andrea:
Thank you, Larry. Thank you for having me. It was a pleasure. Great conversation, yeah.