Compute-ready data for AI with Futurum Signal65
Enterprises are discovering that PDFs are for people; machines need something else. As AI systems move from experimentation to production, raw documents limit accuracy, performance, and trust. This session covers a proof-of-concept engagement exploring why unstructured content must be transformed into secure, compute-ready assets to deliver faster insights, predictable performance, and true data sovereignty. This solution enables processing data once to power AI everywhere in your organization. Brian Martin, VP of AI and Data Center Performance at Signal65, presented the work of their AI lab, which is sponsored by Dell Technologies and focuses on real-world impact tests of AI workloads using extensive AI infrastructure, including various Dell XE servers and NVIDIA GPUs. The lab also developed a digital twin to optimize its physical layout, particularly for complex designs such as slab-on-grade data centers with overhead utilities, demonstrating an immediate positive ROI by identifying costly design changes early.
The presentation then transitioned to the critical challenge of preparing data for AI models. Echoing the sentiment that “garbage in, garbage out” becomes “expensive garbage out” with AI infrastructure, Signal65 highlighted how raw, unstructured data, particularly PDFs, hinders AI accuracy, performance, and trust. PDFs are designed for human consumption, not machine processing. Gadget Software addresses this by offering “compute-ready data,” which transforms unstructured content into AI-digestible formats. This process involves semantic decomposition to maintain topic continuity, LLM enrichment to generate useful metadata such as summaries, keywords, sentiment, and sample Q&A pairs, and robust governance and security through unique IDs and lineage tracking. This approach overcomes the limitations of traditional chunking and pure vectorization, which often lose context and attribution, making it difficult to cite sources or enforce security policies.
In a proof of concept using the vast United States Federal Register, Signal65 demonstrated the tangible benefits of this compute-ready data pipeline. The process ensured that all AI responses could be traced back to the original documents, which is crucial for governance and security. Performance testing revealed a significant advantage for local GPU processing (using L40S and RTX Pros) compared to accessing cloud LLM APIs. Local processing delivered remarkably consistent, flat latency during data ingestion and enrichment, in contrast to the spiky, unpredictable latency observed with cloud APIs, regardless of document size. This “write once, read many” approach ensures that once data is processed and enriched, it can be reliably accessed by various AI applications, such as chatbots or BI tools, delivering consistent, attributable results. Furthermore, the prepared data facilitates user interaction by enabling intuitive dashboards that showcase data content and suggest relevant queries, addressing the common user challenge of “what can I ask?”
Presented by Brian Martin, AI & Datacenter Lead, Signal65, Futurum, Group. Recorded live at AI Infrastructure Field Day in Santa Clara on January 29th, 2026. Watch the entire presentation at https://techfieldday.com/appearance/signal-65-presents-at-ai-infrastructure-field-day/ or visit https://techfieldday.com/event/aiifd4/ or https://signal65.com/ for more information.
Transcript
Hi, I am Brian Martin, uh, with Signal 65 VP of ai, uh, data center, uh, AI and Data Center Performance. I wanna talk today a little bit about Signal 65, what we're doing over there, and share some of the results from a recent POC we just did. Um, so as our president likes to say, signal 65 got its name from the concept of attempting to be the signal through the noise.
Uh, how do you make a difference? How do we figure out what's actually true, find ground truth, uh, in the work we do? And that squarely puts us into the lower right of this futurum, uh, circle, which is the assess.
Uh, get our hands on equipment in the lab, test things, uh, for real, see how they behave, and then write it up and talk about it. I have the privilege of working in and with the AI lab in Colorado. Uh, this is sponsored by Dell Technologies, and we have quite a collection of AI infrastructure.
So being here at AI Infrastructure Field Day feels spot on. Uh, and we all have a mission to accelerate innovation, uh, and do real world impact tests of AI workloads and on AI equipment. This is part of the AI Testing Lab.
This is our air cooled section for now. Um, primarily for these demonstrations of solutions technology built around Dell XE 96 eighties, XE 7 7 4 fives, uh, GPUs from our favorite vendors, uh, up through and including the RTX Pro 6,000 Blackwell, uh, and soon later this year, B 200, B 300, as well as MI 3 55 X. And again, I couldn't ask for a better playground.
Uh, in this day and age with ai, uh, being able to run these models solutions locally, uh, and we do this for Dell, with Dell, uh, for customers, uh, and some of our own investigations. The lab itself, uh, has a monitoring infrastructure that we set up, uh, to keep track of utilization. Uh, the system's being used, how heavily the GPUs being used, how do they running, um, utilization is the only one that's red if it's on the left.
So if it's below 50% utilized, it goes red. Uh, everything else goes red when it gets high. Uh, these, as we know, are not cheap components.
This is expensive infrastructure. We want to keep it busy, want to keep it running, doing meaningful work wherever we can. Uh, also very interested in power consumption.
So power cooling. Uh, later this year, we're also gonna add acoustic as we start to compare air cooled servers with liquid cooled servers. Uh, be nice to know how that impacts the acoustic footprint.
Some of these servers with the smaller fans get very loud. Uh, we're clocking some, you know, 120 decibel sound levels, uh, in the AI lab today. Looking forward, uh, one of the things we've heard a bit about today, uh, a couple of times at Infrastructure Field Day is digital twins.
Um, we did a similar thing for the AI lab, and my team constructed a digital twin, uh, in the lab for the lab, uh, which allows us not only to explore the build out and creation, but also allows us to do a little what if modeling, so we can take new servers, put 'em into the rack. Uh, we went all the way to VR with this solution. So Meta Quest headset, uh, fully immersive.
Why did we do that? Um, going from blueprints to models to fully immersive gave us and some of the lab operators the opportunity to feel the space in three dimensions. Um, we only get one chance to design these labs.
Uh, this particular build out is slab on grade overhead water. It's as scary as it sounds. Um, and in conversation with Dell, uh, just within the last 12 months, uh, there has been a huge uptick in questions from customers about slab on grade in the past, it's always been raised floor water under the floor.
Uh, we're really starting to see companies trying to take advantage of existing space. Uh, as we heard earlier, this equipment can be incredibly heavy. Uh, 3000, 4,000 pounds for a rack is not unusual in these AI servers.
So slab on grade is good for that, but we have to be very careful. So now we're looking at routing, water overhead, power overhead, network, cables, overhead. Um, where do you fit that?
How do you fit that? Um, having VR and the ability for the lab techs and the designers to put on the headset, walk through the lab, get on a ladder, can they fit their head over the rack under the pipes to reach the network cables? Uh, we made a few design changes early on based on these models, uh, design changes that would've been very expensive to do later on in the project.
So not only was it fun, entertaining, and enlightening, it was also helpful with a positive ROI almost immediately. Alright, what do we do in the lab? Uh, I mentioned some projects we've got going, uh, with partners, customers, uh, Dell, we published that on the Signal 65 website.
There's a dedicated page for insights from the AI lab. Um, I'm gonna talk about one of the projects we recently completed around data preparation. Um, this seems to dovetail well with talks we've heard earlier today, uh, about the importance of the data we feed into AI models or the results we hope to get back from those models.
So, Brian, you mentioned Dell servers. Do you have, um, Nvidia, DG Xs and those sorts of things, or? Well, we do have access to Nvidia, DGX.
Those are in a different lab. Uh, not this lab. Not this lab.
They, uh, they're much smaller. They fit into the Kentucky lab. Thanks for asking.
Uh, all right. I think I heard this earlier today. You can't just dump huge amounts of raw data into an LLM and expect reasonable results.
Um, our friends at Forward Network shared that and talked about their solution and why this turns out to be the exact same problem Gadget software, uh, is addressing in disengagement. We did, um, so garbage in, garbage out in this case, it's very expensive garbage out when you're running on systems that cost well north of a million dollars, uh, per cluster. And this gap we get, um, between AI capability, the models we have are incredibly powerful.
Um, what we can do with them, given clean data, clean instructions, clean context is very impressive. But when we don't have data that they can consume or they have too much data, they struggle. Why do they struggle?
They struggle for a number of reasons. They're gonna struggle with quality. Um, you know, in a typical rag environment, uh, we start with data ingestion and that starts with chunking.
And often we overlap the chunks we ingest so we don't lose context on the gaps between the data, but we're still doing it blindly. So we don't have continuity through a given thought or topic. Uh, and then at the back end, um, we have a vector db.
We have our data. And when we're trying to retrieve that data, uh, the LLM has to rebuild the context has to rethread those topics together. It has to understand what's related to what, pull that back out of the vector database, reassemble it, generative process.
You're likely to get different answers every time. Not always what we want when we're trying to have strong attribution or if we have strong security requirements going through this process, we discovered something interesting. And that is, you know, really that the documents we're feeding in, uh, even though they're the most commonly uploaded things into the LLMs, they're not the best format for them to make good decisions.
You know, PDFs and the PDF format PDFs are for people. They're the last stage of publication pipeline, you know, and it's really a publication format. Uh, it wasn't meant to be computer consumed.
It's meant to be person consumed. Computers and AI especially need something else, uh, and what we think they need. Uh, and what, what gadget, uh, is offering is this concept of compute ready data or compute ready documents.
How do we give a computer system something a computer system understands? Uh, and for those of you who've been in and around the storage industry for years, like I have, you're gonna hear a very familiar term and we've also heard earlier today, and that's metadata. And metadata can be a lot of different things.
Um, metadata can be about where the data lives, uh, what file system it's in, what container it's in, what its properties are. Metadata can also be about what's in the file itself about the contents. And this is what we're looking at here.
Um, the metadata that we build up, the gadget builds up in this solution is about what's inside. And they do that in four steps, right? They start by decomposing the document, um, and something they call semantic decomposition.
How do they pull apart a document while retaining topics, um, and concepts and holding those together, feeding those into an LLM that takes those topics and enriches that data with useful things. Um, and those useful things are summaries, descriptions, keywords, um, if you configure or derive sentiment or intent, Silverton consultant, you're not doing vectorization of the documents Or not yet. Well, or should I say also?
Yes. Uh, vectorization also happens. Um, and vectorization of the metadata can happen.
But the challenge we found with pure vectorization is you get proximity results. Like this concept is close to that concept. Um, and what gets lost in that process often is attribution.
Attribution or providence. Um, whether it's security or the ability to simply cite your source. Quick question.
Um, when we're talking about compute ready documents, um, there's a, you know, the legal industry has document management systems that are quite large, quite large vaults of documents. Are they working with you in collaboration on how to do some of the de decomposition? Because they also have some AI tooling around document manager decomposition.
I was just curious if they're involved in any way or anybody is Not yet. We expect soon. Okay.
Um, we're actually, I don't know if it's in the slides or not, um, but we're looking at some grants specifically around this topic. Mm-hmm. Um, which we hope they'll Participate in.
Maybe get a DMS involved. Okay. Yeah, definitely.
Um, gadget goes so far as to even create sample q and a pairs based on the data. Uh, like given this data, what's a reasonable question that might be asked and how would that be answered? Uh, that can be a starting template.
Uh, if anyone's done prompt engineering, one of the parts of, of a prompt that are often helpful are examples. And this starts with some readymade examples. Uh, what kind of documents are they doing this for?
Ah, fantastic. Hold that thought if you would. I wasn't sure where to put that slide, but we do have a use case, a very specific use case with one of the world's largest publishers.
Um, and then as I mentioned, sort of governance and security is a key point. So as it goes through this decomposition composition process and enrichment, they get unique IDs, uh, that help us preserve, um, the lineage. Alright, but what does that look like?
Uh, it looks like a machine. It's a conveyor belt data comes in. Um, interestingly enough, in both the use case we're going to look at and a large portion of the publishing industry, uh, the data phase immediately before conversion to PDF is actually semi-structured data, often XML or something similar.
Um, being able to tap in at that point in the stream helps text parsing, um, and assembling that. But we also have charts, tables on the, the next project we're looking at is basically finding a large volume of data that's literally just scanned documents that are brought in as images. So there's zero text content to start with.
That's when the vision models will really kick in to help, uh, process those, uh, on ingest. The other interesting thing is when this is created, so there's this concept of write once read many for this data ingestion, bringing the data in, validating the data. So there's a, this is part of the pipeline, making sure that the data we brought in accurately reflects the source data.
Once we've got that finished, we basically right, lock it, now it's got an id, it's persisted. And subsequent subsequent processes, whether they're chat bots or agents or um, you know, power BI as an example, BI tools, uh, can access that same data, have the same responses, which bring back for a set of keywords, a set of topics brings back, um, cited references. And then the LLM can do what it does best, which is embellish that for people return.
You can tell me to wait till later on this one too. It's fine. But I'm looking at this, um, this pipeline here, the ingestion, it's data.
So does it have to be a Word doc, a spreadsheet, an email? Can it be text? I mean, like what is, what are the data?
It can, it, it can be anything. Okay. It can be any, anything.
Anything an LLM can interpret. It can be. Okay.
So anything from a picture doesn't have to be in a specific language. Um, what we've done in this use case is, uh, predominantly English text with pictures, uh, for the starting point. So the process of summary description, topic specific keyword, sample q and a, sentiment, intent, topic, boundaries, segmentation is done through LLMs or is that something that Gadget has created themselves?
Or, or what That is done that is an agentic process similar to what we saw with forward networks in the previous presentation. That is a pile of code, the agent that Gadget has written in communication with an LLM to process it. Great question.
The question I have, I just wanna add some context. So I'm thinking about like, use cases of how people would use this and, um, how people create data mm-hmm. Is so sporadic, right?
So I'm thinking of even if you have like a, a, a document that's been all the comments from editing it down, some of those comments are sometimes really valuable informant and, and good things to know, especially if you're gonna ingest it into this type of thing. Mm-hmm. So I was wondering how much, how much they thought about how people actually create data.
Mm-hmm. And how does it, and did that go into consideration of how to make this pipeline? Absolutely.
Okay. So those, those comments, I, I think of those as, you know, comments in the margins. Yeah.
Right. We write in our books, BLE things, highlight things, um, and if those are captured electronically, 'cause you took notes, that's one way. These documents that are photos a lot, a lot of them do have scribbling on them.
Okay. That gets captured as well. Um, and the vision models are getting extremely good these days.
So they can tell when something is in the margins, what it belongs to, uh, and then how it relates to a keyword or a topic. Okay. Cool.
Right. Here's the answer to your question. So the use case we looked at was the United States Federal Register.
This is an, an enormous pile of data, uh, that gets generated daily. Um, they set a record in 2024, over 107,000 pages generated. This also includes final rules, uh, which include a lot of dense legal text, uh, which has to be parsed carefully and has to be referenced accurately.
You're trying to make decisions on this. Uh, so this is a large amount of data, um, in this case in a relatively friendly structure coming in, but with pretty dense, um, content and requirements around it. And at the end of the day, um, the threshold we set is that the tools will have the ability to trace all responses.
Like anything can be cited back to a document. And that becomes crucial when you think about, uh, governance and security. Uh, if you have some documents coming in with certain security restrictions or, um, you know, you can only be viewed by certain people, certain jobs can look at it, then you can have the ability in a system that has this metadata to tag those, to filter those, to block those from going out so they don't come back.
Uh, whereas in a vector database that often gets lost. I'd go so far to say always, but that won't be true tomorrow. gov or whatever that you can access the federal registry?
Yes. By asking, Well, our system is not running there. This system is running there.
Anyone can go to that site and Document. Oh, I understand. But you had, you've got more semantics, summarization, description, keywords, all that other stuff.
Correct. That's Not part of this system. Not part of this system.
Where's your system? This is an overlay That look in the notes, uh, in the video when it gets posted. Um, it, it is from time to time made, made publicly available, um, for this.
So getting to the fun performancy stuff of this. Um, so in the lab, one of the first things we did was compare the solution running in the cloud or running locally, but accessing LLMs in the cloud. Uh, so we want access to the best models to do this work sometimes.
Um, but getting from on-prem to off-prem and back again, uh, managing this data, uh, results in pretty spiky latency. So this is a, a graph of hundreds of thousands of data points, um, run across each, uh, or the entire processing of a month's worth of data from last year. And we brought in each document that comes in, you know, generates, there are some dots around zero, um, generates a little to over 6,000, what are called artifacts and artifact is anything that gets generated from the system, a keyword, a q and a pair, a summary, a topic.
Um, some of the larger documents generate a lot of those, uh, and with pretty predictable, uh, correlation. Uh, the larger the document, the more you're doing with it, the longer it takes. What surprised us running on-prem, uh, with both, uh, L 40 s GPUs and RTX pros is how flat that other line was consistently spot on.
Uh, doesn't really care, uh, in, in the overall measurement. And this is just the ingestion process. It's not the actual query Process.
This is the ingest, ingestion and enrichment process. Correct. Um, when you do these sorts of comparison, the question I would have is, is the hardware similar in both the cloud environment as well as local or is it different?
Are you using RTX six thousands versus, you know, the cloud might be using L four forties or whatever? So in this particular test, uh, this is cloud API against local GPU. So we don't actually know, um, what the system is inferencing on.
Uh, it could be anything. It could be H 100, could be a 100, could be, yeah, yeah, yeah. H 200.
Uh, we don't know what the background Running be some, you know, some slice of something. Yeah, correct. So shared system problems.
Once we have all, all this data, um, what can we do with it? And one of the things that gadget's founder discovered early on, and he likes to say the, the blank prompt. Like, oftentimes you go to a an LLM screen and you get a blank thing and it says, ask me anything.
And one of the things they discovered in conversations with customers is the first question they ask is, what should I ask? What can I ask? Like, what can this data set tell me, um, am I looking at something about environmental law or am I looking at afterschool activities?
Like, tell me something I can ask. And we don't always know that even in our own companies when we're looking at a data set. So one of the things that they did as a demonstrator, this is a Power BI dashboard, um, with a lot of stuff packed into one screen, uh, analyst's, uh, dream here.
Um, but before any question gets asked, like, this is a breakdown of the data that's in the repository. So this is range of topics or departments in the government or areas of the world that are impacted. Um, and it's meant to show different ways to bring, uh, a customer or a user into working with the data.
So there's a lot of different topics, people, places, uh, impacts that can be looked at and drawn through here. So watch for more updates. Uh, as we do more engagements, again, signal 65 insights from the AI lab.
Uh, this is the digital twin for the liquid cooled portion of the lab, which is being built out as we speak. We'll have an announcement on that when it goes live in the near future as well. And we start to see comparisons between the air cooled and the liquid cooled pieces.
Question, This is the actual representation of the data center? Or is This It is the actual representation. It's a big place.
It's a, it's, yeah. 8,000 square feet I believe. Oh, wow.
Um, yeah. This is the actual design, the actual layout, the plumbing or rack spacing. A couple interesting things about this.
You may notice here, there's no hot aisle, cold aisle. Uh, every system in this is room is room neutral for temperature.