103. Waiting for Data Is Not a Technical Problem Presented by IBM – Tech Field Day Podcast Spotlight
Keeping your AI applications waiting for data is not a technical problem; AI agents can help deliver the right data. This episode of the Tech Field Day podcast features Stephanie Valarezo from IBM, who discusses Agentic data enablement with Karen Lopez, Jim Czuprynski, and Alastair Cooke. Organizations face many critical challenges in providing high-quality, secure, and well-governed data at scale for AI, especially agentic AI applications.
Beyond simply possessing data, the challenges include managing vast, disparate data sources across hybrid environments, ensuring rapid access, and building trust through robust data governance. IBM’s Watsonx is presented as a comprehensive data foundation with pillars in AI, Data, and Governance, offering composable products for data integration, quality, cataloging, and the creation of automated data products. Autonomous agents for repetitive data tasks will allow data professionals to concentrate on higher-value activities such as architecture, policy design, and security. Ultimately, enabling organizations to build scalable, automated, and well-governed data ecosystems for their AI initiatives.
Transcript
It's not just data that's holding you back. Your AI needs good quality, secure data, and really the business units that are going to get benefit need to be able to define how they get access to it. Join me as we talk with the IBM team and learn a little more about Watson X.
Welcome to the Tech Field Day podcast, where we bring together a group of IT technical experts to discuss a single idea about a key concept in the industry. This podcast features a variety of perspectives from members of the Tech Field Day delegate community, and it's often recorded in association with one of our events. Tech Field Day is part of the Futurum Group, and this podcast is also published on our sister company site, Techstrong TV.
On this episode, presented by IBM, we're looking at how IBM's Watson X provides a data foundation for agentic AI implementation. But before we get into that discussion, let's meet who's on the panel today. Hi, I'm Karen Lopez with InfoAdvisors.
I'm a data governance professional with a focus on data security and data quality. com, and I focus on databases of various sorts, mainly Oracle these days, but other databases as well. And so yeah, also data quality and how that relates to accurate and more effective generative and agentic AI.
Hi there, everyone. My name is Stephanie Valarezzo, and I lead product here at IBM, focusing on data integration as well as intelligence, and more broadly, how we can put your data to good use. And I'm Alistair Cook.
I'm the event lead for AI Infrastructure Field Day here at Tech Field Day. And one of the things that we see both at Tech Field Day events but also as we look across the industry and how companies are embracing AI, particularly moving towards agentic AI, one of the big challenges, and when we see two really big challenges, the first one not directly addressing here is a skills gap, although maybe we are addressing a skills gap here with these tools. But the other one is data, feeding enough data to all of the different use cases for AI within the organization.
One of the things I've seen is we see these pilots where an easy use case is chosen with a small amount of data and a really well-defined set of data, because that's a great thing to test. But as you start transitioning towards working in production, you start saying, "Well, we've got 400 different use cases. " Doing this by hand one by one, as we did all through our proof of concept and maybe our early pilots, that doesn't scale to dealing with hundreds of use cases, potentially even thousands, tens of thousands of data sources, let alone the concerns we have around the governance and management of that data, because these agentic AI applications, AI applications in general, are being used as an attack vector against us.
And so being able to manage vast amounts of data, feed that vast amounts of data into our AI applications, is one of the fundamental challenges of getting AI into production. It's one of the things that in the Futurum Research studies, we've seen these are big challenges for people getting data to feed their applications and getting good quality data that remains good quality over time. And I've spent a long time hitting this, but Stephanie, do you want to ground us a little bit in what Watsonx Data is doing to help address these issues?
I think if we start off in talking about what we're hearing from our clients is a couple of different things on the data front. One is speed, that they want to be able to have access to data, and it's no longer acceptable to have to wait weeks in order for a pipeline to be in production. By that time, the decision moment has passed.
So that's one of the things. Secondly, and I know, Karen, you're very in tune to this, is how do you bring trust at scale? So really looking at the governance gap, because whether it's for data access, if we're going to try to respond at speed, then how do we enable governance at scale?
Another big question is how do we manage data that's in quite a varied ecosystem, whether data that's still on-prem, it's in different clouds, it's in different SaaS platforms. All of our clients are not going to consolidate their data in one place, but they need access and to be able to actually understand where information needs to come from and the relationships to be able to act on it. So, with all of these different things that are happening, and I'm sure that all the listeners are facing different problems in these different dimensions, the way that we're responding is providing a composable set of products to be able to meet you where you are.
So whether it is that you want to enable streaming data integration, batch integration, unstructured data integration, if you want to do data quality work, maintain catalogs, or be able to expose data products in a marketplace, not to mention thinking about lakehouses and being able to run different engines to be able to allow folks to query or allow you to access vector databases, et cetera. All of those different areas is where our data platform sits, and it's been built over years, especially with responding to a couple of those points that I brought up at the topAll of this is not one separate place that you should get started. You're starting with, I want to expose data for AI, I want to expose data for my agents.
But how do you actually do that? It starts with all the other core components that I mentioned. So you bring up so many points that I want to respond to, but I'll focus on the high level is, these are things that I really think about, about doing data governance engineering at scale.
And by engineering, I mean not just the process and the methods and the theories, like establishing data stewards and all of that, but how do we manage the metadata? How do we manage catalogs? How do we know our data lineage?
How do we manage our pipelines that are moving things around? Or if we're not moving data around and we're responding to data gravity, that the processes that are going to get access to our data, and the people, can do that in a very short period of time, as you understand. I think about even 15 years ago, the main way we did data governance and data governance engineering was with spreadsheets.
And if we were lucky, we had maybe a data modeling tool or some sort of enterprise architecture tool. But even those were focused on a particular application, not really a particular data system. Yeah, Kieran, that's a really interesting perspective.
Strangely enough, I'm actually working on building out more effective and more accurate metadata for a hackathon/datathon that I'm going to be doing next week. And again, right now, the two data sets we have are about half a million rows. Nothing really surprising, but it spans everything from images to audio, to unstructured data, to good old structured data laying around in tables.
From my side of things, it's really, how do you make sure that you don't accidentally give away company secrets? How do you make sure that this table here that no one really knows what's inside it, actually has formulas for if it's a pharmaceutical company, all the things that, for example, make you unique in the industry. If it's a paint company, I have some folks coming from Sherwin-Williams.
The secret to their whitest white paint. Or whatever taupe is. If that were to get out into the marketplace because someone really didn't understand at an elemental level what's inside that container of data, I foresee rather disastrous consequences for an organization because, as Kieran said, they didn't do the forward work on establishing exactly what's inside these containers.
So are there interesting aspects of your offering that help with that? Because that would be something interesting to hear about. Sure.
I love what you both just shared on how things have evolved and the pace of needing to allow folks to be able to access, in self-service ways, data. The risk always with self-service is that you have ungoverned sprawl, and one of the things that we're really looking at is how you have governance integrated within the practices of connecting to these different data sources, doing the transformations, and then exposing it to certain endpoints, which are then what you shop for in a data product marketplace. That's something that we're really looking at is how we can, and actually bought into our products.
How do we lower that floor so that we can democratize data access, but we don't lose that governance factor, Jim, that you just highlighted? How do we make sure that we don't expose risk? Because you need to have access controls.
You need to have those audit trails. You need to have lineage built in. One of the things that we're really looking at is the federated model.
Not everything is going to be centrally hosted. Now with these different teams that are operating across the business and operating at very high velocity, we need to make sure that we can, and obviously these aren't just humans now, these are agents. We have that governance layer built in so that if those folks want to grab data, then we're checking who is that?
What should they have access to? And that way you don't have this risk exposure. Since you brought up data products, I've always been recently thinking about where are agents going to fit in into that enforcing of data contracts in an automated way, not just like the way we used to do it, is if data was the wrong data or came late or was missing data or was the complete wrong data or too much data.
We had automation that would log a ticket into a ticketing system, and then a human was assigned to it, and then a human had to go investigate. We'll still need some of that, but a lot of it's just very repeatable commodity type things to say, "Oh, it failed because of this. " And just have it solve itself, log it, or escalate it to the right people, not just a regular, just throw it into a ticketing system.
So I'm looking forward to agents helping with data products and data contracts We're super excited for that side, too. We're already jumping into the idea of how do we have autonomous data product creation. And the exciting part of where we're jumping to with our approach is we have these underlying capabilities that can then be exposed for agents to then be able to actually power them and be able to create a data product.
If you have the ability to first understand data sources and understand the relationship between them, look and actually know how to build a data integration pipeline and optimize it for different integration styles. Do I need this as soon as there's data that's changed on different sources to be able to access that change? That's something that you need someone that's specialized or an agent that's specialized.
Not to mention, the lineage and governance. And if you have something that can do that for you, it's just a really exciting frontier. governance.
So governance being one of the key tenets in this. I mean, this is fundamentally one of the reasons that people go to IBM, is having that long, long history of managing critical data and looking after it not just day one, but day two, day 200, day 2,000, day 20,000. This is one of the central themes that I thought was really vital in what I saw of watsonx.
But then there's the challenge. One of the big challenges we see, aside from having access to data, is having access to skills. The data scientists who are building out these collections of data to be consumed by your AI application.
That set of skills is a gating limit for a lot of deployment in agentic AI systems and in general, in more complex integrated systems. It doesn't have to be AI systems. " And that there'll be an agent at the background, both initially setting up what a data scientist would do of setting those relationships, getting you a view into it, but also doing the ongoing.
Hitting Karen's point of, what if some data arrives late? What if one of the other teams that's looking after this data makes a schema change and maybe adds some additional information to it? These are the things that, again, would fall back to a human.
And what I could see in the material for this was that the aim is to have agents doing that. That rather than having a human involved every time there is some small change, there'd be that same logic that was able to assemble together the data for me to feed my application, was also going to essentially maintain how it builds that together. I thought that was a really cool thing to see.
Yeah, if I may. The parts that you just brought up on the skills gap and what we're seeing is there is certainly a part that is skills gap, and if we're talking about lowering the floor, then it means that a business user should be able to create a data integration pipeline leveraging technologies like Apache Flink, and they might not understand the intricacies of distributed systems. They should just be able to create the pipeline.
And that's one place that we're targeting. The other part of this is really around, I think the second point is if we are looking for an agent to be able to do some of this work, then what about the role of the data engineer, the data operations person who was doing that work? And we're really expanding that role.
That role is certainly not going away, because we still have to build and maintain these pipelines. But now, instead of just managing the structured pipelines, there might be needs to expand into supporting AI workloads and looking at the chunking, embedding, vectorizing that may need to happen. Or managing streaming architectures.
So it's taking the folks that we have and expanding their realm of responsibilities while we're looking at also changing the number of folks that are able to do this. And I guess the last part that I really wanted to comment on is there's one part of this that's skills, but the part on architecture, and the question that I would have for any of the listeners is, if your best data engineer or your operations folks left tomorrow, then would the remaining team be able to actually troubleshoot? And if your answer is you wouldn't be able to, then it's not even the skills and talent that you have.
That's something that's core to your architecture. So that's something to also balance as we're looking at skills and the data foundation to be able to jump into these new use cases. That's a really interesting perspective.
In fact, that was one of the things I wanted to ask you about, Stephanie, was all right, so you freed me and Karen from tedium. Bless you. Thank you very much.
What-types of activities would you see us now partaking in? I'm interested to hear maybe from what you're hearing from your customers as to what we really would like our data engineers to do would be to more focus on, I don't know, system reliability, security, there's an idea. Or some things along those lines.
What are you hearing from your field in terms of that? That's a great point, Jim. One of the ways that we're helping our clients and what we're hearing from them is data engineers are focused on these repetitive tasks.
Karen brought it up before. Instead of having to do the repetitive pipeline construction or pipeline modification, such as what's needed for data migrations, we need to have them focusing on architecture, policy design, and having the ways that we can automate some of those tasks. Dependency analysis is another one where, as part of looking at do I want to shift to push down, or do I want to switch to streaming, then can I do this in a more automated way?
On the side of operations, this is where having the metadata across your different jobs, monitoring them, and having observability, all of this can then be used to actually improve operations. And like you said, I'm looking at how can I scale, how do I make sure that my systems are secure? All of those different things are places that our clients are jumping to, and they're really seeing the possibilities of what we can do to help.
So one of the things I'm looking forward to with agents is trying to move the detection of these data contract errors back to the source, the beginning of the process. So things like if someone wants to change a table that's being used in these 12 pipelines and also in this completely other system that isn't in a pipeline, I want an agent to say, "Hey, this is great. Did you know that table has these 23 things that are dependent on it?
" And for me, that means that data engineers, pipeline people will be able to say, "I'm going to use this higher level of thinking, of collaborating, discussing, triaging, prioritizing cost benefit and risking," those sort of things. So that I am tired of, even at a low code environment, creating a copy job that copies all the files from this folder into this lake house. And then I also want it to pay attention to when a new file comes in there and kick off an event and bring that over too.
That is such a basic, simple pattern. And yet, even if I reuse stuff, it's still something I have to do. I'd rather just say, "Do this for me, and put in these three exceptions, and never run it at 7:00 PM," and all those things, and let someone else go into the low code, which will be probably a CLI, to do all this.
That's what I'm looking forward to, is for agents to also help me prevent these data governance, data quality issues, not just detect them. Could not agree more. It's one question that we ask clients is, how much of your data estate is actually being monitored for quality?
And what we all typically hear is our critical tables. And now that we are having more access to those non-critical tables, could not agree more, Karen. This is where we then are able to have a better quality monitoring, a better ability to respond to data drift, and if anomalies are detected, if you then have the pipeline, then to be able to jump in and do the healing of the pipeline or to even have that be a gated process with these kinds of changes.
These are the kinds of things that are not destructive. You can update and switch. For these other gates, this is something that is critical.
We need to have someone review. Even to be able to expose something like that. So you come in and log in in the morning, and you're able to see that kind of a view that helps you to better do the jobs that are needed for not just transforming data or moving data, but really making sure that it's ready for those AI use cases, analytics, et cetera.
So that's a really interesting perspective, Stephanie. Just curious, what about how the model or models that you're using to actually do this, how have they been trained? Because this sounds like a very interesting idea, but it's much more wrapped around, shall we say, operations, procedures, and best practices, or simply processes that back in the old days, if we were offshoring this, we would have someone sit down with someone and do knowledge transfer.
Trying to extract that from the recesses of their brains of how things really work. I'm curious how you accomplish that, and then also is there a threshold of it's got to be at least 85% in terms of it's capturing 85% of those processes or those operations to be considered successful. Is there some sort of, how would you put it, processacquisition metric that says we have X percent confidence in this?
Sure thing. So one of the things that we've really been looking at is, how can we expose these different capabilities we have across integration, intelligence, lakehouse, to be able to allow our clients, wherever they might be, they're using Frontier AI agents, to be able to take advantage of the underlying capabilities. And I'll use a quick soccer example, or football, Hal, is if you're getting ready to do a set play in soccer, then ball comes in, and you have a specific sequence for where it lands, et cetera.
What we're switching to is having every player on the field is able to be in any one place. It's a very flexible approach. In this way, we're able to expose our capabilities to be able to be used then for these agentic loops that our clients are creating.
So in this case, it's how do we make sure that you're able to expose a data integration pipeline to create a governance process that's gated, to create a process for updating those pipelines, and to do this in a way that you can then monitor and see what the agent did and set those thresholds. So really, we're taking advantage of the models that our clients are already using, or to be able to switch back and forth between different models, but we're exposing those capabilities on the back end. That framework or that system is the MCP server, the tools, the skills that are all built on these capabilities, the ecosystem really, if you will, that then allow our clients to be able to create these loops around them.
So, I think the place that we're going to, and it's going to be really interesting, is being able to have repeatable processes here. So if I want to have a governance process, before it was, I have very strict steps that need to be followed. And now, when I jump into having autonomous data product creation or assisted data product creation, then how do I make sure that the same input that a user is providing or an agent is providing gives us the same output?
And that's a really exciting part that we're also doing research and building into our framework, is to be able to have that kind of predictable outcome, or the places where there needs to be decisioning to be able to provide that information for users. So that's our approach, instead of building closed black box loops that our clients would not be able to really break out of or use interoperable and with their Frontier models. Well, as usual, we start a conversation here on the Tech Field Day podcast, and it could go for hours.
In fact, this is how our delegates and often our presenters spend their evenings and breaks between sessions at Tech Field Day events is carrying on these conversations. In fact, even inside our internal Slack, we'll have the same conversations amongst delegates, too. So thank you for joining us today on the Tech Field Day podcast.
But before we go, where can people connect with you and maybe continue this conversation? So people can find me, I'm Datachick on most social media. com, and I'm on LinkedIn with the in /karenlopez.
com. I'm on LinkedIn. Look at this last name.
I'm not hard to find. I'm also available on LinkedIn. Same as me.
Last name is not hard to find. com. We have blogs, hands-on labs, things like that, that allow you to jump into the product.
So do go check those out. And of course, I'm Alistair Cook. You can find me on LinkedIn and on the Tech Field Day website and across a variety of different locations on Futurum.
We will also be having more conversations with IBM here on the Tech Field Day podcast. So keep a watch out for those over the coming months. But thank you very much for listening to this episode of the Tech Field Day podcast.
And if you've enjoyed this insightful discussion, please subscribe on YouTube or your favorite podcast application so you don't miss an episode. Do consider giving us a rating and a very nice review of how much you enjoy these conversations. This particular episode was brought to you by IBM and Tech Field Day, home of the IT experts from across the enterprise and a part of the Futurum Group.
com/podcast or view us on Techstrong TV. Thanks for listening, and we will see you next week on the Tech Field Day podcast.