09×02: Moving Beyond Text for Agentic AI Applications with ApertureData
Our online interactions include audio, video, and sensor data, but most AI applications are still focused on text. This episode of Utilizing Tech considers how we can integrate multimodal data with agentic applications with Vishakha Gupta, founder and CEO of ApertureData, Frederic Van Haren of HighFens, and Stephen Foskett of Tech Field Day. After decades of developing AI models to process spoken word, images, video, and other multimodal data, the ascendance of large language models has largely focused on text. This is changing, as AI applications are increasingly leveraging multimodal data, including text, audio, video, and sensors. Many agentic applications still pass data as structured or unstructured text, but it is possible to use multimedia data as well, for example passing a clip of a video from agent to agent if the system has true multimodal understanding. Enterprise applications are moving beyond text to include voice and video, data in PDFs like charts and diagrams, medical sensors and images, and more.
Guest:
Vishakha Gupta, CEO and Founder, ApertureData
Transcript
Our online interactions include audio, video, and sensor data, but most AI applications are still focused on text. This episode of Utilizing Tech considers how we can integrate multimodal data with Ag agentic applications. With our conversation with Vish Gupta, founder and CEO of Aperture Data, Frederick Van Hern and myself, Steven FoST, welcome to Utilizing Tech, the podcast about emerging technology from Tech Field Day part of the Futurum Group.
This brand new season focuses on the practical applications of ag agentic ai, and related innovations in artificial intelligence. I'm your host, Steven FoST, organizer of the Tech Field Day event series, and host of utilizing Tech now for nine seasons. Joining me this week as my co-host is, uh, Frederick Van Herren, who's been present for a lot of those seasons.
Welcome Frederick to the show. Yeah, thank you. Once again.
I'm, I'm, I'm Mihir as a co-host. So my name is Frederick Van Herren, I'm the founder and CTO of Hyen, which is a HPC and an AI consulting and services company. And Frederick and I have been talking about practical applications for AI for a long time, but one of the things that always sticks in my craw, I'm not sure what a CR is, but something sticks there, is that when people talk about ai, they too often focus only on text.
Basically, it's chatbot or bust. And that's great. And in fact, there's a lot that you can do with text, but text isn't the whole world.
Uh, and Frederick, I mean, your background in ai, you, you, you didn't start with text. No, definitely not. You know, my, as you know, my background is in speech and, and that's the language we use to communicate with people, but we have to understand that people are more visual than learning from text.
So a multimodal approach to ai, it's definitely something we're looking forward in this agentic AI world. Absolutely. And I think that, um, you know, just regular people who are gonna, uh, be sort of wondering why are we, why are we always talking about, you know, documents and books and webpages and stuff?
Why aren't we talking about literally everything we interact with today, which is, um, audio, which is images, which is video documents and so on. And that is why, uh, we have an exciting guest to kick this season off. Uh, today we have, uh, Vish Gupta, uh, founder and CEO of Aperture Data, who is somebody that I spoke with, uh, earlier this year, and they are really focused on multimodal data.
Welcome to the show. Thank you so much, Steven and Frederick, uh, really happy to be here. So tell us a little bit more about your background and yourself and, uh, and what you're focused on.
I'm happy to. So yeah, as you mentioned, um, right now I'm co-founder and CEO of Aperture Data. Uh, prior to that I was at Intel Labs for over seven years, which is where we started working on this problem.
And, uh, you know, just take looking at it, um, from a researcher standpoint and now through the journey at Aperture Data, it's been more like product and user standpoint and businesses standpoint. Um, for my background, um, I have a PhD in computer science from Georgia Tech and a master's from Carnegie Mellon. And I got my undergrad in computer science from Biti in India.
So it's been, uh, you know, one part of life where it's a lot of, you know, computer science, deep research, working, really like, you know, uh, on underlying, uh, uh, systems hypervisors. We were one of the first teams to virtualize Nvidia GPUs back when it wasn't even so, uh, popular in, uh, data centers, uh, to, to be, to be offered in cloud environments. Um, and then now this has been a completely different part of the journey, you know, by sometimes, um, miss coding too.
Uh, but it's a lot more about people than it used to be before. And most importantly, it's a lot about very exciting use cases and applications that have emerged in this last decade. Just as you know, we have witnessed the progress of machine learning from, just from the very basic, like, you know, the very first image nest thing that make it like, oh, now we can actually automatically understand what's in an image to now AI agents trying to like, you know, order our plane tickets for us to plan this perfect vacation.
Well, it does seem though that, um, the progress of ai, you're right, that it started, as Frederick said, it started with, um, with speech processing, uh, for the most part, uh, it, it a lot of the applications early in the utilizing journey. Back when we started this podcast, we talked about, uh, processing images and detecting objects and, and, and processing video and sound and all sorts of other data sources too, um, you know, bio, um, mechanic information from sensors and so on. But it seems like the prevalence of large language models has just crushed any kind of discussion of anything that's not text.
Are, are you seeing that as well? Um, I think a lot of the practical applications, you know, the, the approach I end up seeing for people a lot of the times is, well, they're asked to go use ai. So this is like, you know, you kind of have to, you know, we work with a lot of, um, you know, medium to large enterprises, some startups, and we see this very often.
There is always this like, how can we use ai if, if they're doing it right, it would be more from the perspective of, there is this business problem, can AI help us here and then watch the process? But then sometimes it's like, we gotta get on the AI bandwagon. There is a lot of funding for it, right?
So there is like a spectrum of people. Uh, but the common thing is like, okay, look, uh, especially when you think about larger companies, data is very siloed, right? Especially if they're collecting, if they're going beyond simple tabular data, going beyond text, that means sometimes it'll be like images, videos, audio, they get organized in places that most of the company doesn't know, like few people in some team that will have access to it.
So just the process of bringing things together and then making sure that the models are up to par, and then can you actually take all of that and combine that into, um, a model that can give you the right answers? That's a very, um, intensive process and which is what makes it like, okay, well let's prove the value with text first, uh, and then we'll see. Right?
Uh, unfortunately though, the problem is in some cases text is sufficient. A lot of the times people are like, you know, let's say if they are just, uh, uh, there is an example where you have a lot of PDFs and granted it took some while to start parsing PDFs to the level of, you know, actually understanding tables and images within PDF two. It's not, it looks simple, but it's not always, but like, you know, if you're trying to understand a lot of reports that you do internally and enable a rack chat bot, that's the kind of application that you can pretty easily start off with.
Um, um, and, and, and, you know, so you start, if, if you start seeing the ROI good, but if it's the sort of application where a lot of information was stuck in these other data types and you started just with text, it is quite possible that people arrive at the wrong conclusion that AI doesn't work. And I've seen that through a lot of, uh, you know, even before we got to the whole rag and egen stories that we are like Gentech world that we are in today, even when people were just trying to do like, you know, let's say e-commerce personalized recommendation, people wanted to use visual similarity search to recommend products that look like this. Because, you know, we are very visual people, like you said, uh, we are, uh, we look at something and that's what attracts us, not the text description of the product, right?
Um, and a lot of the times teams would start building it, but they didn't have tools to be able to query and see what their image data sets look like. So they would just train the models, you know, with partial knowledge, not do the best job, and then the recommendations weren't as good as, you know, just based on your friend bought this, and so you should buy this sort of stuff. Um, and so the conclusion used to be, well, AI is not helping us.
Uh, and, and so that's why I kind of warn in terms of like, you know, it's, uh, text is a good start, but there are a lot of cases where you got to bring in the other signals and it looks very daunting because the tooling also has been pretty broken. Um, but that's kind of why we are here. Right.
So you talked a little bit about, uh, multimodal. I mean, it's, it's already difficult enough to build models just for text or audio or video, um, let alone the different data types, right? Videoed notoriously are much bigger, binary, very difficult to analyze.
Uh, text is easy to read, but much smaller than video. So how do you deal with all those different data types and those different models into, into one final model, so to Speak? Um, I think I, I mean, I couldn't speak, so the way I look at multimodal data and what it would take to unlock everything that it offers, I look, uh, at it at, in, in three, three sections, so to speak.
One is the models, what you're referring to. And you know, there are a lot of vision language models now that are doing really well. I mean, I'm recently reading all about like, um, generative on the SOA side, but also in interpreting like the, uh, if you look at some of the Gemini output and stuff like that, they can really do a very good job understanding what's happening in image video.
The second aspect is processing, well, there is no lack of processing today, I might say. Um, I mean, Nvidia really has changed the name of that game. Uh, and there's a lot of other, uh, inference provider providers and, and, um, some more niche companies, uh, coming up in that space.
Uh, and then the third aspect is data management. And I feel like this is where the biggest gap exists. If, if you look through the evolution of machine learning, you know, we used to see all these papers where, oh, you know, we are now able to detect like this tiniest bit of, uh, like a dog face in the image.
And then you went and deployed it in like a medical imaging sort of scenario. And I literally have an example I ran where the brain lobe, uh, like the brain scan was classified as a telephone lobe. Um, so, you know, so those sort of things, you, it, they happen because like data has always been the differentiator.
The better data you train, where the more representative data you give to your model, the better the outcome is. Um, and it remains true to this day, but unfortunately, the solutions are still not there yet because changing data systems is a very involved and very complicated process. Um, or it, it doesn't have to be, but that's how we've always seen it.
Like, you know, I mean, there's so much, and like people think so much in terms of SQL and relational tables. Sometimes I run into people, it's like they won't use anything that doesn't support sql, but that's not the right approach. The thing is what solves your problem if, if AI needs to see all of the data, you need a data foundation that allows you to put all of the data, make it searchable, and make it easy to navigate.
That has always been our driving principle behind this, because if you wanna go, um, and you know, earlier we were talking about going from shallow intelligence to deep intelligence, um, intelligence, we are, uh, you know, like I was saying, the models really are very advanced now, the processing power, you know, it's growing constantly and, and, and accomplishing a lot. If we solve the data problem, if we build the right foundation for data layers instead of still cobbling together a bunch of tools that make you inefficient, that, um, create inconsistencies that don't let to you scale as much as you can, you're still never gonna fully go into deep intelligence territory. Yeah, it, it's, it's so true.
And, and what I'm seeing unfortunately is that a lot of ag agentic applications are still focused on structured textual data. In other words, if we're gonna have this process pass data onto this next process, um, in many cases what it's doing is it's devolving, uh, you know, let's say visual data into a, in, in some cases a large set of texts that describes that visual data in either, uh, freeform text or in, you know, structured, um, data and then passes that to the next thing. Um, is it possible for agents to pass the actual image or some abstraction of that image between, uh, agents in these systems?
It, it absolutely is possible, right? So, um, as we were building, so, you know, as you know, the product that, uh, that my company, um, offers is called adb. It's this unique vector graph hybrid database that we have purpose built for multimodal ai.
So there was always this aspect of, you know, if someone says, I wanna find images similar to this image image, there's of course the vector search angle, so you need embedding generation and things like that, but how are people gonna give you the image? Because they would have to give you the image to put in the embeddings, and then you can do vector search, right? They would have to, um, and let's say you were going even further, you wanted to find video clips that had, let's say, kids playing in it.
You would have to be able to understand components of the videos, um, you know, generate embeddings from those and be able to search through them. Um, but it's not just a vector search part because what ends up happening is, let's say you want clips where, um, uh, kids are playing in the video, right? Video, um, video, maybe the entire video is 30 minute long, and there is only like a two minute section in which the kids are playing A true multimodal AI database should allow you to search, decode the video and go into those two minute part and just transfer that.
Now, why is that important? Because, I mean, imagine videos, like you mentioned earlier, videos are really large. Are you gonna be transferring the whole like, you know, 30 minute long video between different components that need to operate on it, or do you wanna just take the two minute sections that are relevant and pass those along?
So that's like, that there is, there is this whole efficiency angle to it and, you know, not having to wait for hours for something to happen that can only be enabled when you introduce true multimodal understanding in your database. And that's, that's kind of what we did because, um, you can literally say, I want all the video clips in which a person was smoking or not smoking, uh, and I want them returned to me in thumbnail size. This is one query to aperture db.
It does the decoding, it tricks out the parts that are interesting and it, you know, you know, bundles up the clips and sends them to the, uh, to the next stage. Um, and you are not duplicating any of this information because remember, you gotta think about scale. We are gonna operate, we, we are operating on petabytes or maybe even zetabytes of data, right?
Um, and when we represent videos in our database, that is the original video file, but then we very smartly use the graph structure that we have to represent all the regions of interest in it. It can be interesting frames, it can be interesting clips in it. It's the same logic with images.
You know, sometimes, um, you might, your cameras might be really high resolution and capture a very wide angle, but all you care about is that person that's standing on the street. Why are you transferring all those pixels between the different stages? Um, so to your, to your original question in terms of why can't you transfer some of these other data?
Can you transfer these other data type? I think one is the protocol that allows you to define the stuff, and I think there is still some room to improve. Like we had to, um, come up with a different query language to support all of this.
Uh, we actively chose not to implement into like A SQL or a cipher based query language because they were too restrictive in what we were trying to do. Now we have built plugins to make them compatible because a lot of other tooling lives in that world, but, uh, we started out without hindering ourselves and, uh, we had to introduce the exchange, like, you know, okay, this is how you are gonna give us blobs of various types. This is how we are gonna DeMar one blob from the other, and the metadata we return is gonna tell you what the rest of it means and things like that.
So, um, we are able to do it. Um, and, you know, we, um, work with PDFs, audio images, videos, and, um, it, it, it works great. And of course, in the backend, uh, we've introduced the whole performance and scale and, you know, understanding that these data types are different.
You have, you know, more parallelism requirements, there is less dependency among, you know, individual parts of the data. So there's all that stuff that goes into the architecture to make it high performance and efficient. And then there is in the protocol to, to enable like to, to define that language.
So it's, it's very much possible to do it Right. Yeah, data movement has always been a problem, and as, as, as people collect more data, I think the data movement by itself will get worse and worse. So how do, how do people interact with your platform?
Do you integrate with frameworks or orchestrators, or how do, how do people use and consume the platform? Yeah, I mean, you know, we started out with a database. No one wants to really think about a database.
It needs to be hidden behind stuff. Um, and so yeah, we have, um, uh, you know, originally when we started, because we were looking at a lot of, um, training and inference sort of use cases, we integrated with PyTorch sense of flow, vertex, ai, these sort of, uh, frameworks. Um, then we, you know, with the RAG in the rag world, basically we introduced Lang Chain LAMA index integrations, and now we are looking, looking into agent memory, uh, frameworks to integrate with them.
Um, we also do, you know, so ADB has grown from just being a database into this entire platform where, um, you cannot just manage and search the data. But, you know, we have introduced workflows to make it easy to upload data to, uh, generate embeddings to extract information. I mean, we have a workflow that, you know, you give a URL and outcomes are at chatbot.
Uh, you don't have to worry about segmentation mo like, you know, embedding models and things like that. Um, so, and, and we have these various, like, you know, MCP server plugins, SQL server plugins, so that has grown, uh, now into a platform so people can interact with ADB directly on our cloud platform. They can use Community Edition, we can do a VPC deployment.
Um, but we also have our own ui. And I'm actually really pleased recently with all the developments that have happened into our UI because, uh, you can literally go to, you know, one of the tabs in the UI type a text question, and if you've like, you know, ingested the data and generated embeddings, it'll show you, okay, these are the images that match, these are the PDFs that match, these are the videos that match and all of this on one interface. And it, there is still so much to, to, um, improve there.
Yeah. So the, the, the, the workflows or the plugins are, are kind of starter kits and I guess for, for, for the, for the consumers. So can you talk a little bit how the platform works with, uh, enabling autonomous and semi-autonomous AI agents?
Right. So, um, there are different ways that you can go about it. Like, you know, a lot of the agents behind the scenes when they want to interact with data, they basically might, you know, just do vector search queries and then, you know, implement their own LLM like feed, like do the semantic search and feed things to the L LMS and then generate the responses.
Um, so that's like very fundamental way, which, you know, you can just use the vector search support we have. You can enhance that with graph rag sort of, you know, like rag improvements to start including the knowledge you've contained in the graph. But where we are seeing this go is essentially introducing this memory interface, uh, because if you look at, you know, the memory frameworks, there are, there's the component that actually takes, um, user log user questions and extracts preferences and, and, you know, relevant personas and stuff.
But underneath it ends up storing this in either vector databases or a combination of vector and graph database or just simple text logs. And ADB is perfect for storing all of that stuff. I mean, the throughput and latency we offer in terms of updates and queries, it's phenomenal.
Um, and so it makes a really, you know, good foundation. And so now the, the thing we are working on now is like, okay, what's that memory layer, right? Like, you know, we start by integrate, we, we'll basically integrate with some of the frameworks already out there.
Um, so the agents can really like, make use of the memory and scale, um, through what Aperture DB offers. So I love this talk of, you know, moving beyond text. I mean, that's the, the, the premise here at the beginning.
Um, but, um, I wonder if you could help us with some in or examples or some ideas about moving beyond video too, because of course, multimodal data, it doesn't just mean video, it means all sorts of data types that are, you know, all varied. So what other data types beyond audio, video, video and obviously text are, uh, people looking at with agentic applications, and what are some of the use cases for that? Um, well, you know, I mean, uh, Stephen, as we talked about, a lot of people are still on text.
We are not even at the audio video stage yet, but, you know, so there is a lot going on with voice. So there's a lot of audio information. I think people have realized there's a lot they can even like, you know, even before getting to voice and, and videos, um, there's a lot going on with PDF because they, I mean, you know, if you, um, we, we, we all create so many reports, right?
And we like to put tables to summarize, we put charts there, um, we have these pie charts and you know, sometimes we put pictures to show the way, like, you know, how our architecture looks. So there's a lot that goes in and parsing PDFs in itself is, uh, and, and extracting information to start making sense and, you know, at what bound, um, parts do you, uh, how do you segment it and things like that. All of that stuff involves a lot of, um, work.
So I've been seeing, um, like people who have managed to go beyond text, a lot of the times it's like, uh, actually this is the kind of progression. You start text vector search, right? Then you start realizing, well, you know, your text is giving you some more relationships about things, and so can we connect and start, you know, utilizing the relations around it?
So it naturally kind of progresses into well graph sort of notion, can we build a knowledge graph? Can we use that information to improve the responses? Then it moves into like PDFs and that auto automatically gets into like parsing images and stuff.
Um, of course voice AI companies. I, I think Frederick, you would know a lot more on this one. They are, they are starting to get, uh, you know, voice in my understanding started with like, let's convert this audio into the transcript, and again, go back to the vector search where I think there is an increased understanding around like, Hey, if you did that, you lose the emotions, you lose the, you know, background information.
And that's sometimes really important. Um, so you go beyond that. If you get two videos, there are some cases, especially in, um, medical imaging sort of cases, you know, where the scans, um, like, you know, nowadays a lot of the CT scans or ultrasounds can be pretty like, uh, 3D formats that can be, um, the neural scans are in a different format.
Uh, so when you go into more, uh, specific, like more domain specific use cases, then the file formats start to be different. So that is, um, you know, can you understand, um, the medical imaging file formats and start and, and enable the medical co-pilot sort of use cases, right? Because I mean, patient information is naturally multimodal.
Uh, then there is, uh, satellite imaging sort of use cases like, you know, what can you gain? And that can feed into, you know, traffic sort of things, or it can feed into agriculture sort of things. But, um, something that encapsulates the GI s formats and, you know, understands the different layering, like a satellite with different resolutions, how do you align pictures from all of those?
So there are different formats for that. And, um, I think there's a lot of development needed on that, on those applications from the model side too. So I think there is that we are gonna see those things come, uh, you know, that'll be more like more dedicated companies first even figuring out the models to operate on these sort of images.
I mean, in the past when we looked at medical imaging, um, formats, we essentially would slice them up. Like, DICOM is like a series of, uh, p and g files, the usual image, uh, format. So we would slice it out because DICOM itself contain too much.
So there's like, you know, um, but that's when you become very domain specific. Yeah, I I'm glad that you brought up, uh, medical, because I think that that's definitely an area where we're gonna see a lot of development in this, but also as, as you talked about a lot of geographical data, um, I was talking to somebody who's working on drone technology and they are working with everything from, uh, you know, GPS data streams to topographic information like you mentioned to, uh, you know, real time feeds, uh, from sensors, and all of these things have to be integrated and localized and plotted together. It was a really interesting conversation a little bit beyond me, but, um, but I could understand the challenges because there, you know, it's not just video, it's not just maps, it's not just text, it's all of these things as well as lidar and radar, right?
And, you know, cameras and, and all of this had to be integrated. So I, I think that increasingly that's what, what the challenge is gonna be is how do we integrate all of this data in a way that a, um, an AI agent can understand and act on without just overwhelming it with data. I, I think you bring up a great point and, and you know, like anything in, in, in AI right now, it's a two part thing, you know, is there's the model it needs to start having an understanding of it.
Um, and you know, there are that the multimodal models are definitely, uh, you know, um, advancing rapidly, and there is that data part. So, you know, one of the unique aspects, the why did we bring in a graph into picture originally it wasn't because we were thinking there's gonna be all this knowledge graph use cases and things like that. We brought it in because it gave us a good way to represent relationships, and it was flexible to let us represent whatever data type we wanted to represent in it.
So in our same graph structure, and we use a property graph structure for that reason, instead of the, um, RDF um, graphs that come in that just like, you know, sub we don't do the subject predicate object representation, we do the full, like the, if there is a representation for people, it'll be like a person node in the graph, you know, name, last name and all that stuff. But it'll be, it can be very easily connected to another node that's a picture of that person, or that's connected to like video clips of that person and all of these special data types videos. Um, you know, we can introduce lidars documents, all these things have representation in the graphs.
You can go from one type to another, and in the same query you can be saying, I want all of the various data things associated with Stephen and Frederick together, like whatever they appeared together, whether there was a text description, whether there was an event, whether there was, uh, you know, recordings, it, it can go, it can use the power of graph reversal to get there. Um, so that's why we kind of, you know, originally started with the graph and, you know, now you can basically represent a lot of, uh, application information in it too. Yeah, I think one of the problems too is that, uh, not only is there a large amount of data, but the, the amount of metadata associated with the data is also getting more complex.
Right. So you talked a little bit, a little bit about the medical and geographical, I mean, the amount of metadata surrounding it is, is creating an additional, uh, problem in the complexity of the models. So, so, so one of the questions I, I had for you was, how do you see Agen AI evolve in the next 12 to 18 months?
I mean, if you look at MCP servers, they're less than maybe around a year old. It's going so fast. What's, what's your vision for AgTech AI in the next 12 to 18 months?
I think there's gonna be a lot more focus on what does it mean to get agents in production. Um, you know, we've built a lot of toy agents, we've built a lot of, uh, like, you know, agents that are starting to do some serious work. Uh, but I think especially in, uh, larger companies, you know, now it's time to go from POCs into production, which means really answer all these questions.
So all that we discussed, you know, how does, how does it get them maximum ROI you have to start thinking about your stack. Like, are you gonna do a framework way? What framework is the best?
What sort of models give you the least amount of hallucination and get you the most distance in terms of, uh, you know, your particular use case? So like, you know, we work a lot in retail and e-commerce, and that is personalized recommendations. Sometimes it, that doesn't require you to be a hundred percent precise.
You know, you're recommending products, you're telling them what you can buy. It's okay if like one of the products you recommended doesn't exactly fall in that umbrella, but we also work with some medical co-pilot use cases, and there it becomes very important that you do not hallucinate. So the guardrails become really important.
So there'll be a lot more increased understanding in terms of, okay, for the vertical that you are in, um, what are you okay accepting and what are you not? And then what does it mean if you wanna go in production, what are all the data types you're gonna have to involve? What teams have to come together to put this information?
What are the guardrails that are gonna be, how are we gonna evaluate? How are we gonna observe and monitor this stuff? How are we gonna capture user preferences at, at scale without disturbing their experience?
Um, I feel like there's gonna be a lot more, uh, you know, uh, focused and organized efforts. And so the tool like, you know, platforms like ours, uh, become really, uh, key in, in making that happen. Um, I do wish though there is also some effort around ping compute.
We've been throwing so much power, and you can see these numbers about, you know, the electricity consumption for AI applications as like literally been drying reservoirs in places because of cooling. Um, I really hope there is some effort around that too, to reduce the energy consumption. You know, it's interesting.
I was just gonna say, it's almost like people need some kind of special database that can handle all this multimodal data and maybe a platform that could bring it all together. Um, yeah, it, it is, uh, I, I think what people need to know is they need to know that such a, that such technology exists and that it is possible to bring together various data types and with AI applications and that, you know, people are working on this because I wonder how many people are just, you know, sort of dismissing it outta hand and saying like, we just can't handle this right, or we don't know how to handle this. Um, so I, I guess, um, what do you see happening next, uh, from the industry overall in terms of integrating multimodal data with, uh, agentic ai?
I think it's gonna, I, I think it's gonna increase at a much more rapid pace, uh, with the, I mean, you know, there is at anytime the big, uh, big companies start talking so heavily about it, you know, they start talking like six to nine months early because they're trying to build up hype around it. But if you went to, um, uh, Nvidia GTC earlier this year, or Google next, or like, you know, reinvent late last year, multimodal was already the thing and agents were already the thing and it was naturally like multimodal AI agents, right? Um, but of course the practicality follows a little bit behind in all of this.
So, um, yeah, so, uh, I, I think we'll see a lot faster adoption, especially like, you know, we are in, we are in production, so, you know, people can really unlock the data part and the moment you unlock data, um, the computer's ready. Alright, well thank you so much for this. It's been a, it's been a very thought provoking as was, you know, our previous conversation and I hope that our listeners are starting to say, wait a second.
Maybe it's not, you know, just about text and just about structured data and, you know, passing JSON between, you know, agents and things like that. Maybe it's, maybe it's more than that. And hopefully that's the sort of thing that can come from this season of, uh, utilizing tech where we're gonna be talking to a bunch of folks who are doing some really cool things with AI agents.
Um, before we go, um, please, uh, let us know where can we connect with you, where can our listeners connect with you? Where can they learn more and where can they con continue the conversation? Yeah, so I am very active on LinkedIn, so please connect with me on LinkedIn.
I suppose you'll share the profile, um, as part of the description. io or um, docs aperture data io. Uh, and I would say give it a try.
io. Um, you can try out the database, you can try out our various workflows that make it really easy to ingest existing, you know, data examples, run some embeddings, try out the ui, everything is there. And if you are concerned about, uh, privacy because you know, you work at a company that won't let you send data to a SaaS tool, then we also have, um, free community edition on Docker hub and you can definitely try all the database features, uh, through that as well.
And we would really like to grow our community. We have a Slack channel, uh, and we really, you know, amplify people who build and contribute, uh, to the set of applications that can help end users. Um, so for sure, looking forward to such contributions and more multimodal agents built on top Of Apple tob.
Yeah, I can't wait to see what people build. Um, and Frederick, how about you? Yeah, I'm also active on LinkedIn.
com websites. And you will see both of us at AI Field Day, uh, which is coming up real soon here at the end of October. Uh, we're pretty excited to, uh, be bringing together a cool group of companies, uh, talking about various, uh, elements and aspects of ai, some of whom you will hear about on this episode, or this, I'm sorry, on this season of, of utilizing tech and, uh, hopefully some of whom, uh, we will connect with further.
Uh, if you are excited about AI and, uh, agentic AI and, and where this is all going, uh, do check out the Tech Field Day website. com. Uh, that's the website, um, tech Strong AI is our media site.
And also, uh, we're gonna be launching another podcast, a weekly podcast focus on AI as well. So keep an eye out for that. So thank you so much for joining us and listening to this episode of Utilizing Tech.
Uh, you'll find this podcast in your favorite podcast application, just search for utilizing tech. You'll also find us on YouTube. If you enjoyed this discussion, we'd love to hear from you.
Please give us a rating. Please give us a review. Uh, please subscribe.
Uh, this podcast is brought to you by Tech Field Day, which is part of the Futurum Group. com. Or connect with us on X Twitter, uh, blue sky, mastodon, or, uh, yeah, LinkedIn.
Uh, you can look for utilizing tech. Thanks for listening, and we will see you next week. Okay.