Unleash AI with HPE and Kamiwaza
At AI Field Day 7, Robin Braun presented HPE’s “Unleash AI” initiative, emphasizing their collaborative and outcomes-based approach to bringing AI to practical use. Braun was joined by Kamiwaza CEO Luke Norris, in all three sessions. HPE launched the Unleash AI program in early 2024 with the goal of curating a robust partner ecosystem, offering customers end-to-end solutions that are pre-validated on HPE infrastructure and easy to deploy via the channel. They highlighted the importance of converting AI hype into real solutions by working closely with ISVs, creating relevant demos, marketing collateral, and training resources to make AI more accessible and actionable for enterprises. The program’s global scope and diverse use cases, from Vision AI to agentic AI, demonstrate HPE’s commitment to addressing the real-world needs of customers across various industries.
A key focus of the presentation was the Agentic Smart City AI use case in partnership with Kamiwaza and the town of Vail, Colorado. This initiative is a practical example of how municipalities can solve operational challenges using AI. By working with Vail, HPE and Kamiwaza developed several use cases, including improving ADA Section 508 web compliance through AI agents that identify and remediate accessibility issues, saving time and avoiding costly manual web redevelopment. This project broke down data silos and enabled interdepartmental collaboration without requiring cloud connectivity, as everything runs securely on HPE infrastructure. The result was not only a technically sound solution but also a model for how public agencies can adopt AI incrementally without excessive risk.
Kamiwaza’s agentic AI platform, demonstrated during the session, operates as a full-stack orchestration engine capable of connecting to and processing data across distributed environments using various hardware and AI models. Whether running on-premises or at the edge, it brings compute to the data while abstracting the underlying hardware, which enhances performance, scalability, and flexibility. The system incorporates advanced features like ReBAC, an enhanced role and attribute-based access control framework, and ephemeral sessions to enforce security and privacy rigorously. It enables enterprises, including government entities, to be “unbound” by token-based AI billing models and instead focus on fixed-cost, outcome-based deployments. These capabilities have already shown transformational potential in environments like Vail and attracted significant interest from large global enterprises.
Recorded live in Santa Clara, CA as part of AI Field Day 7 on October 29, 2025. Watch the entire presentation at https://techfieldday.com/appearance/hpe-presents-at-ai-field-day-7/ or visit https://www.hpe.com/us/en/private-cloud-ai.html or https://TechFieldDay.com/events/aifd7/ for more information.
Transcript
Delighted to be here today, and I'm joined by Luke Norris from Ka Waza as well. And we're gonna be talking today about, uh, well, obviously ai, but really starting with Unleash ai, giving a little bit of background on what we've been doing from a program perspective. Be able to talk about some of the announcements we just did at GTC, but really focusing on the partnership that we have with Kami Waza and some announcements that we just made, um, actually yesterday about Smart City ai.
And bringing what, you know, to tie into what Steve was talking about, bringing, making Agen AI and AI practical and tactical. It's kinda my favorite way of looking at it is there's a bunch of hype, there's a bunch of, you know, promise and kind of the gleam in an engineer's eye, but how do we start to make it practical and tactical for our customers to actually get benefit from it as opposed to just stand in awe of it. So, um, so I joined HPE at the beginning of March and, uh, and I lead the Unleash AI program.
So I wanted to explain a little bit about the program, 'cause we've been moving very quickly. And, uh, so we announced 28 new partners at the, um, at the end of June at Discover. Um, it was a, it was a quick blip in the, in the keynote, however, wanted to explain a little bit more behind it other than the pretty picture of logos.
And it really is working out in curating an innovation ecosystem instead of partners that, that are well respected in the industry and that have interesting, um, capabilities to help create that last, that last mile of the solution. Because in the end, our customers are looking for, and our partners are looking for the entire solution. How do you bring together that total picture, that total value prop that somebody can use?
It's not just talking about the blinky lights on the server. And don't, don't get me wrong, we all love talking about blinky lights on servers and how fast they go, but in the end, what people are actually wanting to talk about what people wanna solve are the challenges that they have on a day-to-day basis. So we bring in and work together with our ISVs, uh, to be able to, um, to create a, a really bi-directional partnership.
And it's something I'm incredibly passionate about. My background is, I've been in the channel, I've also worked with AI ISVs, and now sitting here from the HPE perspective, how do I make this partnership meaningful to everybody involved? The first goal is that it is a channel first partnership.
So while we'll go through, we'll pre validate the technologies on HPE infrastructure to ensure that the rollout is very simple for our customers. We work first with the channel to be able to, uh, to be able to package this together, to may be able to make it simple for our customers to be able to access these particular use cases and to help drive activation and go to market in the channel to bring awareness. 'cause there's so much noise in ai.
How do we help people know, Hey, we can help you solve your 5 0 8 compliance challenge. Or, Hey, your need to improve your fire detection. How can we help do that in a very straightforward manner?
So as part of the HPE Unleash AI program, we work with our partners to be able to create collateral, create demos, do all of those things as though we were going to sell the software ourselves. We do enable our field. Uh, however, we also work to enable and grow all of this out in the channel and work with our channel partners to take it to market.
What it allows us to do by doing all of that, it, it creates a much better end customer experience. And in the end, the goal and, and the real passion that we have around it is how do we create that faster time to value and not falling into that canyon of despair that people have somewhere between pilot to enterprise scale. Um, and I drive marketing crazy.
I like talk for like a whole long time on a slide, and then I'm like, oh yeah, and here's the other three slides they gave me. Um, so the, uh, so the Unleash GI ecosystem, this is the ecosystem that was announced at the end of June. We also, um, announced Nexus MD last week at Health, was it last week?
Last week at Health. Um, and you'll see, continue to see announcements, discover Barcelona's coming. Ooh, maybe there'll be announcements there too.
Uh, so this ecosystem will continue to grow. The big thing is we're not just looking to create logos on a wall, but to truly activate and to work together with individual ISVs as well as a collection of ISVs like we'll be talking about today, um, in bringing together, um, value add solutions for our partners and, uh, and customers. Um, so this is a wide range of different use cases of different, um, of different types of ai.
And that was all very intentional. Uh, it is not just a North America program, it is a global program. You will see additional, um, European and, um, apac, uh, partners being brought into this as well as a number of these can work across different areas.
Uh, but you'll see everything on here from vision AI to agen AI and everything in between. And, uh, so it's, it's was very intentionally done to start to open that aperture of how can we help our customers solve the challenges they have today. So this is just some of the kind of example use cases just to kind of set the table and then really dive into where we're working with kaza.
So when you start to look at financial services, being able to look at that compliance, scaling risk score, modeling, process automation, again, bringing very specific use cases because a lot of our, a lot of our partners can do a whole bunch of things, but it's, but the, the way I explain it is that if you can do everything, then in the end, a seller doesn't actually hear that and they hear that you do nothing. So how do we help give them that context and that ability to start a conversation so that they can, they can correctly help connect the dots between that total end solution and the, um, and the, um, ISV and then all of the supporting infrastructure. And yes, security we do know is under is being acquired by Veeam, but we're not, we won't change that until that actually closes in mid-December.
The, um, before somebody decided to point that out. The, um, healthcare and life sciences, that is a passion area for me. Doing some really fun things there and adding some, add some new additions, manufacturing and retail.
Obviously a, um, large space where there's a ton of capability and opportunity to bring forth. Um, how do you take off those mundane tasks or how do you process huge amounts of information at once given the dynamic and ever-changing environment that we live in. You know, are tariffs going to be this or this today?
What is the potential impact on supply chain? There's going to be a hurricane here, is there an impact on your supply chain? All of those sort of things.
How are you able to do that, particularly when you start to look across the breadth and depth of people's supply chain as well as product mixes, legal, public sector. But where we're gonna focus today is on smart city. And so yesterday we announced the HP eent Smart City solution.
Yes, I've been practicing that. Um, the, uh, thank you, I read it wrong five times. Um, uh, but very, uh, very proud of what we've been able to do in a very collaborative approach, working with the town of Vail, we'll go into that a little bit more.
But looking across different initial use cases, again, with the notion of kaas as that one Ag agentic AI platform that will continue to plug not only these use cases into, but additional use cases as we go forward. When I looked at this approach and when I was thinking about the smart city, um, everything is always about vision. You know, we, um, we, we think about, uh, you know, you always see those big knocks where they can do all of the, all of the vision analysis and they have an iot, but I'm unlike cities do more than have cameras there.
There's a lot that when you think about a smart city, it's not just a smart department, but smart by smart department, how do we help the city itself cross communicate and not just create additional silos of a, you know, additional silos. Now reinforced by ai, they already have silos of data. They already have silos of operation.
How do we not have AI reinforce that, but instead help to break those down and bring those communications across the city as necessary or as helpful, as well as look at those mundane tasks. And we'll, and we'll double click on this further, that would actually free up the city to deliver more citizen services, free up budget to be able to focus on the delivery of what the city's therefore, which isn't necessarily having somebody manually go and look up records. So we'll dive into this a little bit more, but really delighted by the partnership with Ka Kaza Pro, HAWC VIO and Black Shark Pro Hawc.
And, and vio, for those who aren't familiar with them, do Vision ai Pro Hawc is focused on doing video restoration. I don't say enhancement because they do not use generative ai. They're actual restoration techniques are, are, um, able to be brought into a court of law.
So they're not actually manipulating the image, they're just ex exploiting the different, um, pixels that are actually captured in the image that we don't initially see. Um, and you'll be able to see examples of that, but then actually brings in vision analytics, um, as a vision AI platform, black shark ai lets us extend that out into geospatial. Again, some interesting different use cases and how you use these different types of information to be able to make different, um, decisions and in and, um, considerations.
And then all backed and brought together by Kami Waza. So Luke, I'll pass it to you. I wanna try to fit in on the camera here a little bit.
Um, I guess we're just gonna juggle back and forth. Are you seeing customers, uh, adopt your use cases? I'm sorry, what Are you seeing customers this for?
Robin, are you seeing customers? I'm Ray Lucchese, silver King Consulting. Thank you.
Are you seeing customers adopting some of your use cases? Uh, yes. Um, and we can speak to a little bit of that.
We're gonna go through use case by use case, but in particular, one of the things we just did in partnership actually with Kami Waza, and I guess I can yes, You back in The, this has to be hysterical on livestream. The, um, so the, so we just released in September the, uh, around 5 0 8 compliance. 5 0 8 compliance.
For those who aren't familiar with it is for if you have, if you receive in federal funds, uh, you have certain requirements in making any information that you put out there accessible for people with disabilities. So think about al text on images, being able to read a PDF, things like that. Um, and when you look at it, um, and we'll show a little bit more details on this, I'm stealing, uh, stealing future thunder.
Uh, there's only about like 23% of like federal websites are actually compliant because it goes into a lot more considerations and a lot, and they have different state regulations and a whole bunch of things with it. We'll get into the, what's interesting is putting out that, uh, that offering together, uh, we have had, we're having a number of customer discussions and interest veil has already, um, has already implemented it and is and is absolutely thrilled. The town of Vail, he's from Colorado.
Yes. Yeah. And everyone else would recognize that.
Yes, yes, Yes. The town of Vail, that skiing place, um, the, um, the, um, so, and so what was interesting is that the conversation is a hundred percent on the challenge that they have with being able to disseminate information. Um, like with talking with Russ, uh, Russell Forests, who is the town manager of Vail, he said, you know, people are actually taking information off of websites because they can't figure out how to make it compliant fast enough.
And that's not the idea behind the regulation. The regulation wasn't to remove public information. The regulation was to make it accessible.
But if you're gonna be penalized for not having it accessible, people are figuring out different ways through it. Instead, we've actually been able to, working with Kaza, they created an agent, agent that not only goes through and tells you where it needs to be remediated, but is actually able to do those remediations. We keep a human in the middle to be able to review and then post up to the website.
Uh, but the level of interest has been, um, quite frankly, astounding, uh, the number of meetings, the number of follow ups that we have, because we're not going in and saying, we want to sell you agen. We're going in and helping you not go on a three-year mission with your website to redo it, just to make it compliant. And because before, really the way this has been done is primarily through web developers and consultants.
And, you know, when you go on a multi-month, multi-year mission, we've all dealt with websites and we all know that that is one of the least fun projects on the planet. No disrespect to all web developers and designers. Um, but they can be very difficult and very costly for municipalities and education systems.
And so they're incredibly interested in this discussion. And again, what I love about it is that I'm helping them to solve an issue. Yes, it will take HPE infrastructure to help do it.
We can all, we can install all of it with, um, privately within their own data center, um, connected to their data behind all of their firewalls. Nothing has to be connected to the cloud or to the internet. Um, and so it allows them to do that.
And, but it gives us a business conversation and that we're creating a solution that solves their business need in an economical and viable way. So for the geeks in the room, and I sort of preempted, I put some geek slides in here. So you guys Yes, exactly.
You are that one. Well, I, I, I don't think any is excluded. Yes, exactly.
Uh, um, we, we we're gonna do just a little deep dive on what kaza is, some of the backline tech. So then when we actually show you the, uh, actual use cases, uh, which will be real live demo, well not live demos, but demos that we've recorded, uh, we can go step by step and you can see how sort of it's all connected together. Um, we try to differentiate ourselves effectively off of these four major pillars.
But when you guys think about, uh, KA waza, think about it as an orchestration engine. We're truly trying to sort of delineate ourselves in that we're gonna connect all the data of an enterprise, wherever the data is in any format, it is in these exact location it is, and the exact type it is. And then we're gonna connect all of its security.
We're gonna create an ontology scheme on it, and we're gonna have that flow through multiple log agentic systems, legacy systems, legacy file formats to make all this work. I believe Veil started three months ago, and we have four massive outcome use cases that is transforming that town in less than three months. That's from literally the HPE dropping the server in us, connecting the data, working with the business lines, and actually getting those flows.
So once you connect all this up in this particular way, it just goes incredibly quick. So from a deep dive perspective in, for some of the scowls in the room, let's just talk about what a distributed architecture and mesh is here. First.
It's a full stack solution. It gets installed in a set of docker containers. In the case of, uh, HPE, uh, private cloud ai.
It's, uh, wrapped up in Kubernetes and can be deployed and scaled pretty much in any environment. We call it opinionated, but loosely coupled. You can literally download it, get it up and working within about 15 minutes.
It's about 176 packages for the full enterprise AI experience. Loosely coupled enough if there's already a solution in place or they started a journey or middleware can integrate with say, a vector database that maybe is already in service and solution. Last but not least, we really wanted it to be sort of hardware agnostic.
And what I mean by that is it's gonna work from everything from an Nvidia Blackwell to the RTX 6,000. We're gonna talk a lot about today. It's one of our favorite, uh, chip sets, uh, all the way to Intel processors and other more unique processors in the public cloud.
So it can span everything from the Tanium trip sets, uh, up at AWS uh, uh, T processors to Google, all the way down to HPE gear on-prem and at the edge. And that's very important for us once again, 'cause we wanna access all data, all locations, all formats, all services, and orchestrate that together now where we really differentiate ourselves. Yes, please.
This is Keith. This is Keith from the visor bench. You just mentioned a very, very, very different sets of architectures.
I I don't want to just gloss over that I'm struggling. So when I'm thinking about GPUs, a MD rock, 'EM versus, uh, Nvidia Cuda today, I have to make choices. You're telling me I'm not making a choice.
What am I losing in that, in that abstraction? Well, I wanna say you are making a choice based off of the thermal load, the cost to run it, and the amount of memory you need to actually load the system in there. And we're actually enabling you to have that choice versus being limited to just what capabilities are presented to you.
What I mean by that is the actual software, when it installs, inspects the underlying hardware and it will then deploy the inference engine that's needed to run on that particular chip, set our capability, and then we expose out, uh, uh, two APIs. We have our own standard API, and then we have an open ai, API from an inferencing. So from an app or service perspective, it's always just hitting those standard APIs with the hardware completely abstracted underneath it.
Uh, it will literally run on your Mac laptop all the way up to, like I said, A TPU and Google. Hopefully we'll go more into this. I, I have questions.
So, Um, so ask away. Yeah. So from, from from that hardware delineation, we can now install this software, uh, anywhere and everywhere.
Now, where it really starts to quote unquote delineate even further is the concept of our distributed, uh, uh, inference, I'm sorry, our distributed data engine and our inference mesh. So the inference mesh, think about a data mesh. If you have our stack installed in multiple locations, they connect over, uh, your private enterprise links.
We create our own networking above that via vxlan, typically using Docker Swarm or other technologies. Once again, that's automagical. You just point the two stacks to each other.
They'll negotiate what they can do to build that, uh, VXLAN link once they're connected. Now in this mesh, we can move an inference request the operations of AI based off of about 40 dynamic features. The most important one we redirect off of, and that's everything from like model type, from vision models to language models, to underlying hardware to availability of tokens, et cetera.
But the most important redirect is based off of data and data locality. So when we prep all of the data, when you install our stack and you point it to all the data, the systems of records, the flat files, the object files, we're gonna talk about some cool legacy files in the Vail use cases. When we scan all of that data, we build not only metadata of that data, we build an ontology scheme of that data.
We tie that to the OAuth and SAML authentication of that data. So the true enterprise security structures of that, all of that goes into our local data catalog. And then the local data catalog, when they're connected in this inference, mesh creates a global data catalog.
And through that we tie each inference, request the authentication of it, and then a recursive lookup to what data is gonna be required to answer that inference. And that's live in the stream, in the stack. So when you put all that together, you have this paint by numbers on the right, an app, an LLM enabled service, a user hings our stack in the theoretical cloud in this case, that stack, when it gets the inference request as a recursive lookup to see what data it's gonna need to answer that inference request, where that data resides, and whether it's the security of that data and does the user or the agent that is actually pinging it have access to the location type of data and the access to that data.
So look at this, this sounds magical because let Usually Not a compliment coming from this crop, uh, the magic smoke an infrastructure world, uh, I I know you have history with storage. So as we think about traditional global file systems, we've been trying to solve this data locality problem of being able to move the data before we tried to move the data to the compute or, uh, so you're saying instead of, uh, taking this approach where you're moving this, the data to the compute, which has always been a bad model, we're moving the compute to the data regardless of the under underlying compute. So in my data center, I have big GPUs, I can do all kind of fancy stuff at the edge.
I get a, uh, I, I get a weird vision, I get a weird, uh, video from the vision, uh, inference that's running there. Mm-hmm. Mm-hmm.
Send that to the centralized compute. It says, oh, that's different, do this different inference thing that we haven't established yet. Now run that, even though that's running on cp.
Correct. Okay. Um, also, this is why I don't mean to keep dropping hp, and I'll do that plenty of times.
I mean, HP's ed, thermal compute loads, I think can get up to literally almost 250, 300 degrees ones that we're testing with, uh, the DOW recently out in Hawaii all the way to the DL four eighties with just the RTX 6,000 and air cooled servers all the way up to, you know, the new G 300, uh, and L 70 twos. So being able to have that wide scope with our software running on it all, redirecting where the inference goes and what data's next to it is why we are so fortunate to partner up with a big partner like hp. So let's talk about that as, as part of the challenge.
So I, one of the challenges enterprises have is sizing, especially sizing at the edge. I don't know how much compute to put somewhere. Are you providing telemetry data back to, uh, the folks running the entire distributed mesh to know that, oh, we need to call HPE GreenLake to replace out these, uh, CPU bound systems with LL or RTX thousand, Not L forties or whatever the, whatever the, the latest.
Yeah, the, the latest and greatest. Maybe I have some L forties laying around, I don't know, whatever the, whatever the case, uh, are you providing that telemetry data back so that I can resize my infrastructure is Needed? Ne Next slide, I'll, I'll dive into that.
But our, uh, app garden, which are agents that have heads on it, tied to our tool shed and tied to our model deployment engine, and they make full recommendations on what app you're deploying with, what model you should deploy and do you actually have the hardware and the capability and that, and it's all shown in a nice gooey on the stack can all be also discovered via API. We do get that telemetry from all of our customers except the DO uh, D and, uh, IC community. Uh, so we have a wide swath of deployments across Fortune 500, global two thousands that make those recommendations.
So If I have a, I'm sorry, Ray, one last question. So if I have a process that needs to run on the edge, but I don't have the capacity to actually run it, I'll basically know before I deploy the That's correct. It will actually tell you, uh, because the agent that you're deploying to do that, like a visual language model with an app, it'll tie those two and say the underlying hardware won't support it.
Luke, are you saying you're going to direct an inference based on thermal requirements? We absolutely do. Nice jaw drop.
Yeah. Yeah. That, that was good.
Um, so did we get a camera on right? Yeah, exactly. So, so I mean, just to finish this off here in this sort of, uh, modality, you have that app that service, it pings our stack in the cloud.
The stack in the cloud. Does that recursive lookup of that metadata? Make sure that a, the user or the agent that made the request can actually access the data.
And then where the data is, what it realizes in this particular example is it is it needs data from both the cloud and data at the data center. So it runs the rag process with the data in the cloud. It then sends a, uh, uh, an informed inference request to the stack in the data center.
That stack gets hit. It then runs its rag process with its data locally via its security in the data center. It gets a summary of the answers.
Keep in mind this process could be 10,000 plus data pools. So for every inference, the actual amount of data that needs to be pulled needs to be re-rank, needs to be re-ran, needs to be re pulled, needs to be re-rank to get these very long. Agent ones is typically somewhere in the range of one to a thousand, a one to 10,000.
Like it's a massive data structure, massive data pool. It does all of that locally. And then it sends the summary back to the initial stack that made the request, in this case, the cloud one, it concatenates the two summaries and gives a single answer back to the user that made the request or the agent that made the request.
In doing this, we process entirely locally, but we summarize globally summary results could literally be sent over a dial up modem, a 1920 literally kilobit modem from the 1990s because it's just small amounts of text at that point. And it's just recon, any those rerunning an inference, putting them back together and giving a single answer and doing this. We have transformed the ability for enterprises, as you'll see with Vale once again within three months, to do massive processing and capabilities entirely locally, entirely in their distributed environments as well.
We'll get their answer back where they need it and get it presented out. And we are seeing an uptick that has just been phenomenal, uh, really helped and led by our partners at HP Shunned. Okay.
All right. Uh, Scotton solution, if you're going here, great. But I would love to spend time on all the agent definition.
What are your agents really doing autonomously? What speed bumps have you seen? How have you improved based on speed bumps and things of that nature?
Yeah, Uh, so we have four live demos that, uh, Robin and I will go through, not live, I, I'm sorry they're recorded, but they're actual demos of actual use cases at Veil that we just, uh, showed at GTC. It'd be great when, if we break those down 'cause then there's something meet there versus me ESO typically throwing it out. Two more last slides and, uh, the next slide, we'll finish off some of the questions we got there.
Uh, to do this too, we had to redo the security paradigm literally from the ground up. So we invented something called reback and like everything in security, you need another acronym, right? So, uh, uh, our back are the traditional role-based authentication doesn't work in the agentic world.
You have a user or you have a subagent that typically makes a request to a larger agent. That larger agent now has a much larger amount of data. It can access and security structure.
That agent might talk to another location and another agent there, which might have an even larger amount of data that it can access, et cetera. If at any one of the points you get an inference of data sets that are larger than the initial user request or agent has access to, you can have a much elevated answer from a security posture and procedure. Because we understand all the data of an enterprise, we have the ontology of it, and we have the metadata structure of it.
We can build an attribute based structure along with the relationship based structure and the role based structure, put it all together and you have reback. But that allows us a user that has only this level of context to talk to an agent that has this level of context to talk to an agent that has this level of context. This agent knows only inferred data that the user should actually have.
So it's not an elevated structure that they get a response back at. Think about it in HIPAA environments, how important that would be because you could cross not only HIPAA zones, but GDPR. Think about it in KYC environments, you know, know your customer in, uh, financial services, um, where you might have a larger customer that has many sub LLCs.
You might have an LLC that's part of that user but shouldn't be able to access other ll C structures and so on and so forth. In the agentic world, this has been a game changer. This has been adopted in the highest levels of the DOW and the IC community based off of us.
And this is being adopted in Fortune five hundreds left and right. I probably talked to at least 75 of the Fortune five hundreds in the last month about this particular feature alone. Not only do we have active deployments with HPE, we have many POCs that are running in the wild right now.
Like many, uh, all based on testing this out and using this. 'cause this is the main thing that's really held the enterprise back data access and security and getting agents to actually work and do the work. Last but not least, oh, Can I back up on please?
Yeah. The reback piece, given the non-deterministic behaviors of large language models, I assume you're doing some deterministic agent tool building, but then they, I assume, I guess I'm just trying to build up like how that actually works logistically, because you're gonna have different kinds of datas. They're gonna require different kinds of policies.
You're gonna have to read those in and, and be able to apply these kind of attribute based policies. Sounds like by hand or no, Fully our ology is fully automagically built by subsets of models. What Worries me is the non-deterministic aspects, It's actually very deterministic.
So from a a vectorization standpoint, we actually use like BM models, which are just static models for actually building, uh, those representations. They're not LLMs, so not, and then that actually creates, uh, the vectorization also in any of the security constructs. We actually have the LLMs write code and it's executed code every time for understanding it.
That's exactly it. So those are entirely deterministic all the way through. Further, just because I, I love hip hop, that's Keith or somebody knows, uh, we have a product called Run Runback Turbo.
We actually run anything over 50 times and we can discard any non-deterministic calls outta that. So out of a 50 inference in a row, we could actually say, look, these 13 are variants, but we have 37 that were accurate. And it actually then goes with the 37 as the next one forward.
So effectively, because we're non token bound, we're using big HPE servers with big Nvidia cards, we can literally run tens of thousands of inferences even if needed to, and get that deterministic piece down to a almost a near science. Like we're looking at 98 to 99% on all processing we do in the enterprise. So Luke, this is a, a pain point for me again, Keith from the advisor bench.
This idea, I think one of the things from, and I'm going to be a little controversial here as if this is new Nvidia and Jensen talks about the token economy, but the reality, if anyone has used a, a code agent or anything that needs this type of, uh, recursive type of analysis, PTOs I talked to have no interest in paying for tokens or whatever. They just don't want to pay for tokens. They want to pay for outcomes.
What have been your conversations when you, when you tell this process to customers who are considering on premises solutions versus I think even when I talked to the big cloud providers, none of them are saying that customers are coming to them asking them per tokens per, but you know, how many, how much am I paying per token? Um, so, uh, marketing wise, we're talking about the, the being un token bound or right, or the unbound economy. And the reason you need that, and you, I believe I can pause a couple of the vail use cases.
It's just to build the ontology of a decent city, uh, of Vail. It's somewhere in the range of like 35 to 70 billion tokens. And that ontology probably needs to be reconstructed monthly, quarterly off of change data, new data, et cetera.
So that would mean somebody holding up one of those open AI plaques from the dev day, like every month about, you know, I'm one of the top 10 users of OpenAI type scenario when you can be unbound by tokens. So things you can do are just astronomically different than what goes on there. In each one of the use cases, we actually try to have our agents now actually show the amount of tokens they're processing 'cause it's in the millions and billions to accomplish these things.
It's a, it's a concept that the code sort of generators, the cursors and open ais, they don't show that because it's actually pretty small tokens, input tokens and very small output tokens. What we're accomplishing with our production use cases, once again, being unbound with these RTXX thousands and these, uh, B three hundreds is amazing because you can't throw so much horsepower at it. And that's what the businesses don't care.
Gimme a fixed cost and a real outcome and they're signing up for it. Give me a variable outcome and a variable cost. And it's never been an enterprise adoption Except for the cloud.
Except That, Except no, I actually totally, I totally, I totally don't agree with that. The concept of the cloud is you can go in and burst up and burst down, but I have yet, and I mean yet to work with a Fortune 500 that is doing more than about 5% variability in the cloud. It is fixed workload, fixed workload, fixed Workload.
Well, they just reduced that market to the Fortune 500 where sure, the variability of a few million dollars here and there is, uh, real important in the departments, but not so important. And, and, and in ai that's almost entirely why you're starting to see massive repatriation because these GPUs are running 24 7, 7 days a week and they're not paying for the cost structures in the cloud. And neo clouds like core weave that are adopting four enterprise workloads and use cases are pulling those workloads outta your general top four clouds like constantly.
Yeah, I, I agree that repatriation day has finally come after Yeah. You've been talking about it for Exactly. No, it's That.
So no, with your, um, and control of data, how are you handling that or how are you also sandboxing, um, CPS tool calling and even tool registry altogether? Yep. So, uh, tool registry goes into, uh, our orchestration platform and what we call a tool shed just to have a sort of cute little vernacular.
The tools are tied to agents, the agents are in our app garden. Those two then tie to the metadata and the authentication structure of the actual inference itself from the agent. So you have an agent that can only call on certain tools.
Those tools have to be in the authentication structure of the user and the data, the tool calls, whether it's SAP, et cetera, has to fall within that reback authentication. So once you've wrappered all of that and connected all of that data, that's what allows these agents and use cases just so You, let's use your exact same HIPAA workflow. You have a user making a request and the workflow ends up hitting an agent that makes a tool call, um, the response data to the tool.
How are you sanitizing that and cleaning it up before it gets dumped back into the tokens? User makes a request to the agent, the agent then verifies that the tool it's gonna call is gonna be able to access data that the user actually has access to. 'cause we understand the metadata of even the SAP that the tool would actually be able to call our structured data systems like sql.
If it doesn't, it actually gets denied all the way through. Does That have to be built into the tool wrapper? It's built into the app wrapper, which is where the agent and inferencing happens, and you get a list of the tools and the tool shed that it can actually access and what metadata it actually accesses.
And that flow sort of pulls it all together. Uh, we have it really beautifully shown in the gui. Um, but really it's all API calls 'cause you wanna structurally understand all of this through our system.
Yeah, okay. Kind of building on that. So because that's always been the, the challenge with a lot of tools is not everything allows the ability to do like assume rule, um, type actions where it's like, I've got, I i I need to be able to do this as this particular user, uh, so that I, I get restricted by the RBAs that are built into that application for those things that don't have that ability to, you have kind of some additional overlays that you're able to, to offer to be able to, um, help build that functionality out.
Yeah, so that's the reback piece. So the actual metadata construct of the data, we can understand the relationship of it. So maybe it's financial data, but maybe it's financial data based off of data centers.
You can then flag that so that A CIO would actually have access to that or the CIO shouldn't access the actual underlying p and l of say the organization. And, and that's a very, very keen thing. So, uh, within agents there's the concept of the mosaic effect we call it out here.
You can have one copy of data that's non attributed to another, uh, set of data, but if you infer those two together, you can actually have a correlation. And what we're trying to do is understand that and stop that if it's needed type scenario. And that's a sticky problem that's been around for about 70 years.
Agents make it 10 times more of like an issue. 'cause now you have this inference, this godlike capability kaza, and we want to actually be able to sort of separate that and break that out. So this is our first major approach for it.
We believe we're the only ones tackling this. We have patents around it, but more importantly we have massive installations on this. How does that roll into, um, your token and context management?
Uh, don't know about to, you mean tokens as in like data manage? Yeah, context management. Uh, so once again, all of that, uh, uh, uh, structure would flow back into the actual input tokens of that agent.
Uh, it would either be CV C at that point and then that's goes into the output, but they won't even get to the agent if it doesn't have access. If the original user or agent that made the request doesn't have access at the metadata or relationship level, Do you have a workflow that's actively managing the context or is it just part of the, the message bus that's happening? You can do both.
So you can do the responsive data in or you could actually have another subagent do guard railing and strip out things like PHI data. Um, But you have to do that yourself. Put, uh, we have the workflow there.
You would've to tell that subagent what guardrails you wanna put in place. Yeah. Uh, last but not least, 'cause I wanna move quickly.
We have a lot to cover. It's just this idea of ephemeral sessions as well. Um, you guys ask some great questions, but if you have a session, so you've copied data that you can have in there.
Once again, you're, you're now inferencing data that maybe is related, not related and you're getting new answers, you're getting new data created, you're getting new concepts within that session. We use deprecating uh, tokens via JWTS so that the second that session's gone, you can't rebuild what's within the session. So that's a flag we turn on for healthcare environments, financial service environments.
Think now you could actually have a boardroom agent assisting the board and you wouldn't be able to pull that back up in context post board meeting. It actually allows you to have these sort of private inferencing sessions with data that you should have allowed. And that's other unique feature we had to build for, uh, the government agencies, uh, and large financial institutions and Fortune five hundreds.
So quickly moving on, this is the eye chart. This is actually, it's all the questions you guys asked. This is what the actual stack looks like.
Um, I know it's an eye chart on purpose 'cause it is a large orchestration engine, but as Keith said, like how do we actually keep inspecting and doing this? When I say it's that full stack delivery, this is it. The important part to think about on this, we're trying to lay this orchestration in, lay the data, lay the security, lay the context so then anybody can create agents on top of it.
And all of our Unleashed AI partners can connect to us via our API, our SDK, our MCP. We actually do provide our own MCP server and client so that you are connecting to us, we're connected to the data, we're connected to the security, we're the inferencing point and we can present back out. And that's how we can start to pull the unleashed AA partners, HPE and other people capabilities.
And I kid you not, I wanna keep saying this, Vail started from concept three months ago to full production for incredible use cases. And that isn't, uh, undue timeline. We have Fortune 500, some, 150 years old that now have 20 or 30 use cases with us within the first year.
It is once you get the hard part done, getting agents to work across the enterprise is the fun part. So Keith Townsend from the advisor venture, I hear the speed, but what's on the soft side? What's the disconnect?
A city government that moved slow politically was able to do this in three months. What was the, what was the, the catalyst that helped them understand that they'd be willing to take this level of risk? Because this is a risky project?
Project? I think we, we took a really different approach rather than coming in and saying, you know, like, hi, here's our agents. Which ones would you like?
Um, we instead came in and talked to them like a partner and started with and, and this was, I thought Russ uh, had some great description of this. Russ Was the city manager, The city manager. Um, as we were talking at GTC, we came in and did initial education about different types of AI that we thought would be applicable for their scenario.
Things like vision, things like working with Kami Waza. But we did education and started talking to them about what we could do while also asking them about their challenges. So we did some initial education and then did a two day workshop where we, where we really went in deep on kind of what's keeping them up at night and what's keeping them from maybe going to kind of the next level kind of from each department as well as continuing that education and being able to bubble up kind of these four initial use cases.
And then we've been working with them collaboratively as we've built out these use cases to ensure that they were hitting the mark for what they needed and aligned to not only their brand and approach around citizen services, but also they have this fantastic view of everyone really as a customer with, with so many visitors that come, that come in, you know, Vail has over 2 million visitors a year when you think about during kind of four months of great powder, uh, you know that there, that there's a significant influx of population into the Vail Valley. So how do they, how do they deal with all of those different changes? So it's been a very collaborative approach and, and I thought, um, Russ had some great conversation around, they started as kind of, they needed education on ai.
There, there were skeptics to, to education needed. And that by going through this process holistically, rather than coming in and saying, here we can do this, this, and this, which do you want? But instead doing it more collaboratively that, uh, that it really has changed and, and that the growth has been significant in the last three months.
I think also because we've been able to deliver several of these use cases and that people who started maybe as skeptics but came along on the journey are now asking the, well what's next? What if we type of questions? And so we're scheduling already scheduling kind of that next workshop to go through and say, we've done these excited for this.
What are this, what ifs and what's next now that we're starting to open the aperture of what can you actually do with the ai? Um, and, and where can you see benefit? You know, it's funny, this is everything we have play out here.
I keep rolling back to, it's the same stuff we face in network automation. It's the same stuff we faced in, um, infrastructure automation, you know, and and for us it was the new stuff is the code and now the new stuff is the ai. Yeah.
But every conversation always seems to go back to those same type of points of sitting down talking listly, what problems do we need you to need solved? Yeah. But what's really interesting about this is that the business is, the line of business is making this decision of automation is, is difficult enough to, to convince it, to adopt automation, but then to go to the line of business, the line of business decides to adopt automation, whether you call it AI or whatever it is, it, it is scary and it's impressive to you folks that you were able to convince.
I would, I would see the government to do this. I, I agree. And I think the difference is automation is a subset of what we're talking about here.
There's a lot more wrapped around this, including the stuff that Keith just, um, but I think from a brass tax perspective, just like with automation and now wanting to let little robots do things autonomously, we really need to think very, very thoroughly about building trust. Mm-hmm. Mm-hmm.
And how do we get from, you know, I'm, you know, I can issue a command and it gives me a result to releasing, you know, an agent, um, to go do things autonomously and then maybe releasing it once and letting it keep doing it over and over and over again. And like that's, that's where the skeptics are. You know, they hear our agents do blah and what they really wanna know is yeah, what are they really doing and how do I get from what's in the marketing to what are they really?
And I think, and I think it'll help as we go through the use cases because that makes it practical and tactical Yeah. Where everybody can see it and it becomes much more tangible rather than kind of the interpretive dance approach. Yeah.
Um, That's also your, your point there. You know, it, it also goes into, you don't just throw stuff out there on trusting, you know, you just don't make production changes on Friday. What you know, mean down, have write down.
Right. Do what society would like to have a review. But my, my thought on that is that you send the agent, you send the robot, you send the human whatever is gonna do the task needs to have adequate training and he has to have adequate skillset be able to accomplish it and they need to approve that they've been able to do that task.
Yep. It's no different for human as agents. And I think once we can get past that and actually do it in a training environment, then it's easier to release.
And I think also one of the things is that we chose very specific use cases so that we weren't coming in and trying to boil the ocean. I like to say we were actually coming in and boiling mud puddle by mud puddle. Now after a while you put enough mud puddles together and you get a lake or you get an ocean.
But like we didn't go in and try to solve everything day one. We came in very specifically with the audit, with looking at, um, housing deeds, with looking at 5 0 8 compliance everywhere in there. We do actually keep a human in the loop, but I think because we were very specific about what we were doing, now we can build on that to go into the what if.
So I just wanna point out, we announced this yesterday, so October 28th. So you guys are the second purview of this. Uh, but these are live, these are in production and Veil is very proud of what they've accomplished and what we've accomplished with them.