Why Uniqueness Defines the Future of Enterprise GenAl with Articul8
At AI Field Day 7, Dr. Arun Subramaniyan, Founder and CEO of Articul8, presented a compelling vision of the future of enterprise generative AI (GenAI) centered on hyper-personalization and domain specificity. The talk addressed the evolution of AI accessibility, highlighting how general-purpose models have become commoditized, making it easier than ever to generate content but often sacrificing depth and context. Dr. Subramaniyan emphasized the limitations of relying purely on generic large language models (LLMs) for enterprise tasks, especially in specialized domains like healthcare, manufacturing, cybersecurity, and energy. Instead, Articul8 advocates for tailored solutions built using domain-specific models that can factor in both tacit knowledge and contextual nuances that general models may overlook or misinterpret.
Articul8’s technological strategy involves developing proprietary models specifically trained or fine-tuned with datasets relevant to particular industries. These models range in size from a few billion to several hundred billion parameters. They are not designed to be general-purpose but to perform specific tasks within domains such as thermodynamics, aerodynamics, and system diagnostics. Articul8 also builds task-specific models, such as those focused on table understanding in spreadsheets or parsing complex PDFs. These models are augmented with multi-model orchestration capabilities and metadata tagging to construct automated knowledge graphs, enabling richer semantic understanding without moving raw data—only metadata is transferred. This approach allows enterprises to retain control over data security, including air-gapped deployments, while benefiting from AI-powered insights.
Articul8’s platform architecture supports both on-premises and SaaS-based models, enabling flexibility in deployment. The platform can ingest unstructured data, apply semantic reasoning, and, using a combination of models and tools, generate agents and squads of agents that autonomously execute “missions” within enterprise workflows. This layered AI model architecture culminates in creating “digital twins,” which are dynamically updated representations of business systems for simulation and high-fidelity analysis. Demonstrating measurable performance gains, Articul8 benchmarked its energy domain model against state-of-the-art general-purpose models like LLaMA 3, GPT-4, and others, showing superior results across multiple task dimensions. The company maintains control of intellectual property in models unless customer data is involved in training, in which case the IP resides with the customer. This approach ensures scalability while maintaining rigorous standards for contextual accuracy, enterprise integration, and domain relevance.
Recorded live in Santa Clara, California on October 30, 2025 as part of AI Field Day 7. Watch the entire presentation at https://techfieldday.com/appearance/articul8-presents-at-ai-field-day-7/ or visit https://TechFieldDay.com/event/aifd7/ or https://www.Articul8.ai for more information.
Transcript
I'm Arun Subramanian, founder and CEO of Articulate. Um, and we are thrilled to be here to discuss what we are building and where we think the world is going from the perspective of generative ai. And, um, we thought the, the focus of the session today, in addition to just talking about what we are building also in terms of where we think enterprises should go, right, and hope we can have a, a great discussion.
Love to get your points of view as well. And we do have some, uh, demos planned today. We have, uh, coming up and, uh, REO coming up in terms of talking about the product.
And hopefully the demo got smile on us today. So, um, the world we see today is on the left side, which is basically everybody getting excited about AI and getting access to tools at different points in time. And today, pretty much every enterprise, every um, say technical person on the planet, if they want to get access, they have access to more or less the same things.
And imagine how quickly that has happened to the point from, oh, these models are so dangerous that they cannot be released. That was not that long ago to pretty much all models getting released to features being so unique that they were all behind proprietary firewalls to pretty much any feature from a general purpose model that if you want to use today, you kinda get access for almost free. I mean, the latest announcement was OpenAI announcing that they're going to open up their entire pro subscription to anybody in India for free for a year.
Just imagine that just in terms of customer acquisition, like the bar for what you can do with these systems have gone way up and the bar to actually get access has gone way down, which means everybody generally produces content that's significantly better than what they used to produce from a volume perspective just a few months ago. And I wanna stress the volume thing because the quality, unfortunately, uh, it's gonna sound strange coming from an AI enthusiast has gone way down. Okay?
And the reason for that, and you can see, uh, any number of different articles coming out about AI slop and what's going on around the world. It's not even AI slop. I have seen cases where people have put in a lot of effort in generating content, creating data, collecting data, and then using an AI system to summarize it, which proceeded to destroy all the work they did in collecting the data.
Because it doesn't understand the context in which it needs to analyze the data and produce the report that actually takes away all of the meaning they put into the data collection process. That's what I mean by the world being very gray today and getting out of that is really what gets us out of bed, which is, today's world may look like this, but where the world is going is hyper-personalization, meaning it needs to be unique to what you are doing as an expert at that point in time. Meaning if I'm asking for a summary of something, if I'm sitting in Santa Clara versus I'm sitting in Japan, the context of what I'm asking is significantly different.
The systems know where I am, the systems kind of know why I am asking something. I would like them to actually respond differently based on what they've learned from me in a safe way. That's what we mean by personalization.
And also it's the uniqueness of that particular person, right? So why would we go to any, um, say a particular expert that we trust because we care, we value their opinion, even if they are 180 degrees opposed to what we think we have to be doing, we would still value their opinion. That's the uniqueness we want to continue going forward.
And that uniqueness can only be infused if we work with the systems individually as individuals. It's not easy to do today, and our job is to actually make systems that are very easy to do, right? That's really where we would like to go.
But the other notion also is the whole notion of general purpose models versus domain specific models. I like to motivate that concept. Maybe I'm, uh, preaching to the choir here, but this was not really a concept that was, um, uh, well accepted not that long ago, which is if you have a complicated question, okay, the question say happens to be about brain surgery, you give a textbook about brain surgery to a high school student and tell them that the answer is in this textbook, find it for me.
Most high school students would be able to find the answer. They may even give you the right answer. You ask the same question to a brain surgeon.
The brain surgeon in front of you picks up the same book, just refers it before giving you an answer. Which answer would you prefer? Even if the answer was identical?
The answer to any, uh, room that I've asked that question is very obvious. Obviously the brain surgeon, because they have context, the experience, the rightness of the answer doesn't matter as much as the context in which it is actually delivered to you. That's really what we mean by domain specific models versus general purpose models, right?
And we actually, um, believe that general purpose models are very much necessary. We are greatly thankful for all of the models that have come out. In fact, we are even more thankful about the research they've put out there.
If you look at the, the LAMA two papers, the LAMA three paper and the four paper, the paper is way more valuable than the actual model, which, right? However they are necessary but not sufficient to get you to your outcomes. And that's really where we come into play.
We're going after domain specific data that may even be in the public domain, but processing the domain specific data in a way to build models that understands the domain is what we mean by domain specific models. And of course, getting into proprietary data sets the values obvious, but we believe that going after use cases that are truly differentiated for that particular domain requires you to move to the right side table. Stake use cases are on the left side.
What I mean by that is if you go into an enterprise that does manufacturing, you make them more productive by giving them a better HR process or a better finance process, or a better customer service process, absolutely Spread. Luc from Silverton Consulting, um, I, you know, a lot of the contextual, uh, information that's provided to the LMS today through rags and other solutions to manage the context. Are you, are you talking about something different Than that?
I'm talking about more than that, right? The reason for that is when you give a context to something like a rag system, you already know what context to pass. That's number one.
The second one is just because you sent the context doesn't mean that the answer would be any different compared to just reading the context only in the context you sent. I'll give you a very specific example, right? So for example, if you are, uh, an automotive designer and you have to design a particular part, now the part has a particular specification.
So you can give the specification as a context to a general purpose model and say, does this meet that specification? The answer is never zero or one. It might say that these are the specs that are met, these are the specs that are not met.
However, even though the specs are met, this is not something that is general practice in that particular domain. It is not a context that that model knows. And as an expert, you would not tell that to an expert.
It's like saying you walk into, uh, a car dealership and say, I want to buy a car. You are not going to say, I wanna buy a car with four wheels. It's as obvious as it sounds in many, many domains.
A lot of these, uh, like obvious nuances that are taken for granted are not known to a generalist, right? It's that context that is unsaid that the domain specific models need to learn, right? And in many cases, of course, if you train the same kinds of models with the that kind of tacit knowledge, they would potentially learn that.
But it's not just the models, it's multiple models working together in sync. That's what we mean by context. But I'll give you many examples of where, uh, like just providing the right context to a general model would fail.
And we have seen many, many examples like that. I use the design example because it's very obvious to give a small box. The box is actually much bigger than that.
So I'm trying to be sure I can conceptualize this. So I have a process, I have a rack process that, that has two steps. First step is to collect data.
Mm-hmm. What I want to do at that point is use a small model mm-hmm. To collect and triage the data.
The next step is to, uh, analyze that I might use a larger and different model, uh, for that part of my process. So what you're saying is, what the difficulty is, is having consistent context across, over the Yes. And we are even saying something.
So when you're processing data, right? So take the, because both of you use RAG as an example, right? So when you're taking any data sets that are large, breaking it down in order to be accessed later on, you need some models to analyze the data that you're processing.
Now, if you take an image and say, create a description of image, you use a general purpose model to create a description versus a domain specific model to create a description, you get very different descriptions. And what you store dictates what is your retrieval accuracy and what you retrieve dictates what is your generation accuracy. I'll give you a specific description for that.
This is, uh, an image from the energy domain. And if somebody just presents this image to you without the one on the right side, it's not that easy to figure out what that image is to a generalist. It just happens to be from a high voltage tension cable testing system.
Now we have an energy model that's built on energy production and transmission data sets that understands that the model never saw this image either. So that this is a, an actual test image. However, it saw many similar things.
It understands the context very well. It obviously gives the description that is very well, right? And this test, so, So you're not, you're not talking about fine tuning No, No, no.
You're talking about retraining or training a specific LLM like level model for an, for an energy domain. Is that what we're talking about here? More or less?
More or less. So I'm giving you a sense of, uh, like what is about to come, because the notion of fine tuning versus continuous pre-training versus pre-training, all of those things have to be dealt with from the perspective of how much data you have, how much uniqueness does that data bring, and what is the complexity of the task you're going after, right? This example, for example, you can do with fine tuning if that's all you're wanting to do.
But if you want to teach the model to learn more context, then you need to make sure that it forgets all the general stuff. Because, I'm sorry, Keith Townsend from the advisory bench. So for us geeks in the room to help me drive this home the point, let's say that, you know, generally speaking, I can go to a, one of the big public models and say, help me with this error code that I'm getting.
And it's pretty good at that. Good. What I can't do is say that, uh, uh, if I'm in, uh, a healthcare environment and I'm not a mainline, uh, care piece, and there's unique limitations to that error in my environment, yes, I need a much more sophisticated, capable Model, Sophisticated capability.
'cause it knows that I can't just change, I can't simply just go and change the security rule to solve the problem. I have to work around that security rule that's unique to my environment. That's Correct.
But also the notion of thinking that any one model would satisfy all the conditions in a complex environment mm-hmm. Is, Is just not possible. That's really why the system that we built is a system from the ground up that understands that you need multiple models to work together.
And the reason I I'm showing you this example is the model that we are comparing against is the state of the art model. Now, if I had told that model that, hey, this is an image from the energy domain, it would never say it's a bowling alley. All I had to do was, this is an image from the energy domain.
Now gimme a description. It may not say exactly what the domain specific model is saying, it won't be accurate from the perspective of being specific, but it won't be as far off as saying it's a bowling alley. That's why we all are okay with about 80% of the thing being there because you know what you need to prompt, but what happens when you don't know what you need to prompt?
Mm-hmm. That's really where the domain specific models really shine. So down deep, um, the, essentially within a rag pipeline or within a data set, when you vectorize it or, um, you bring it into the system, you're, you're attaching additional data to it.
You're attaching metadata. And that metadata, if, if that's where y'all are going, that metadata is where you now have intelligence, okay? On the information.
So let me throw out an example. You have, um, a phone tree of, um, people to contact on any given day for tech support. Um, and those are the on-call rotation.
And so you, you dump that into a vector database and it just has a list of names. And so then you have to parse that information and go somewhere else to find that name and la da la. But if you just simply attach their badge ID number as metadata inside your own vector database, you can then have intelligent tools pull in more enrichment data, real time to do that for you.
But It goes both ways. Imagine the same metadata being corrupted by a description like that. Now you think that your system has information about bowling alleys, but it's actually about energy systems.
So this, we see this all the time when we, I love your content example in the beginning. You've, you've spent all this time collecting data, creating report, and you want, uh, AI to, to refactor it into a summarization, and then it adds some detail that you never mentioned. And now it becomes part of the corpus of data moving forward.
And it was ne it injected, uh, and injected. Yeah. And it's, it is, that is nothing nefarious about any of these processes, right?
No. The person had good intent. They want to generate a quality report.
They might have missed three sentences in a 20 page report. They thought that they gave all the data. However, the specificity with which we need to check all these things actually has gone up, not down.
In fact, if you ask me if I spent 30 minutes writing a draft before today, I probably spent an hour, but the final draft is way deeper, lot more research, lot more content. It actually is a simulated draft because I've asked multiple times, what would the audience actually think about this multiple times and going back and forth. Many times, the longer it takes, the more quality that comes out.
However, I actually need to understand that if I thought, okay, I gave this data, let me generate a, a draft in 10 minutes, that's a dangerous game, right? So that's really the world that all of these things would go. And by the way, all of us would expect that from everybody else.
How many times have you gotten emails that were very obviously something that the sender did not read? I mean, I've gotten emails where the last sentence is, do you want me to do more than is coffee paste? Right?
Not ill intent, it's just that they're trying to do something. But then now imagine if I take the same email, send it through another AI system and say, how do I need to respond? And then that chain keeps going back and forth so our bars will go up not down, right?
And, um, uh, in interest of time, what we've built is something that understood this upfront. We knew that we had to build multiple models. We knew that we had to build domain specific models to get to the outcome.
We've been doing this now for well over three years, and we really built in intelligence layer that understands all of the different systems we put in place and can at run time decide what to pull in. So all the things that you see here are the outcomes that our, our platform can actually do. You, um, and we'll walk you through the individual aspects of it as well, but the highest level, we have complex unstructured information coming in.
We need to be able to perceive meaning from the complex unstructured information that's coming in. And the reason why I'm using these words very specifically is without the user or the end user telling us what's in the information we need to be able to perceive the first step. After that, we apply their context on top of it.
Now, I'll give you an example. As a taxonomy, you walk into, say, an energy company. They have taxonomies of different equipment.
They have different systems, they have different end users. They have, they would like us to ingest data sets based on the taxonomy, but if we apply it before we understand what's in the raw data sets, we would actually miss the connections that they don't really have a, a, a manual sense into. And the second thing is we also won't be able to find any of the mis tagging information that's gone on over decades.
So that's really what happens in the first step. So everything you think about in terms of semantic understanding and data perception comes from there. So the words that you hear from us may be very similar to what a lot of other companies might use.
For example, all of this ends up in a massive knowledge graph. And the knowledge graph is built automatically, meaning the parent child relationships in the graph are not manual input. And that's important because if you have to do that, then graphs become very cumbersome.
So even if you take, say for example, A-A-P-D-F document, when human beings see the PDF document, we automatically build a map in our mind. That map is what we are trying to recreate. Meaning how many images are there, how many tables are there, how much text is connected to what images?
We rarely read A PDF from the top left to the bottom right. We scan, we go to the most interesting thing, we dive out, we go to the other most interesting thing, we dive out, we break it down and then put it back in our heads as a map. That's what the system is designed to do, not just from a PDF perspective, but if you have a database, you have a data warehouse, all of those things.
And the design intent is also that data does not move into our platform for us to go do anything. We only move metadata as you talked about. Very rarely, unless the customer actually wants us to move the data in the raw data stays, stays wherever it's, we only store the metadata that we generate from the raw data.
And the downstream use case is also, if you notice, they're all based on the outcomes that an enterprise would go after. The highest level, of course, is what we call a digital twin. Just in terms of the hierarchy of things.
You start with data. You build models using data, you combine models with tools, you end up with agents, you combine multiple agents with maybe some datasets. You end up with squads of agents.
We use squads of agents to go solve missions and multiple missions put together ends up becoming digital twins. They're a living, breathing thing. You can use it to simulate information.
Um, and it is about saying chatbots are as stable stakes as it gets as useful as they have been. They're usefulness in the enterprise. To get to complex use cases is something that you need to go build on top of.
That's what the, the platform is actually designed to go build. So platform something, a SaaS service science thing that runs in the cloud. You provide the yes data OnPrem to some extent, and you extract metadata and goes out to the cloud and builds these domain specific models in the cloud.
So it's both. So we are an enterprise software platform, so you can, so it's designed to actually run entirely effect the enterprise security perimeter, right? So, and it also runs as a SaaS platform.
So you can consume in both ways, but Uh, just be clear, you're, you're dealing with something, uh, let's call it a pre-trained LLM and you're building on top of that or you're starting from scratch. So we do both. So we've built well over 20 different models.
Some of them are completely started from scratch. Some of them started from pre-trained models. Some of these LMS take hundreds of thousands of dollars, you know, massive data centers and you know, yes.
So it's been an expensive proposition for us, but they do that partly because the outcomes are justified. We have not built any general purpose models. So no general purpose models are something that we've built.
We've only built domain specific models that have been built from the ground up. Yeah. So, So if you're, if you're not trying to solve every problem with a model, then building a for any specific model size, are we talking about 7 billion parameters less?
Like what The, so in the, so we've built everything from, uh, a couple of billion parameters all the way north of a few hundred billion parameter models. So these are vertical specific models. I mean energy, this is, uh, you know, fabric and bakeries.
Absolutely. They're vertical specific models. Yes.
And once you built that sort of model, you can actually apply that model to other customers in that, in Absolutely. So other customers, other domains, even for example, if you take, uh, say thermodynamics or if you take aerodynamics for example, you build that model, it's the same model that works in automotive, the same model that works for aerospace. So it, there's actually a lot of cross pollination between models that actually go through and we'll talk about that as well.
So how many verticals have you modeled? Roughly about seven verticals. We will talk about that.
And in addition to vertical models, we also build task specific models. I'll give you a pet, uh, task of mine, table understanding. So every model has gotten significantly better than understanding tables.
General purpose models will get you about 80, 90% of the way the last 10% is close to impossible. And the reason I say that is take a table with columns that are merged, or cells that are merged, or rows that are merged, take tables where the, um, the units for the numbers in the tables are not in the table. These may seem obvious things for human beings to go solve.
Very hard problems for machines to solve today. And especially if you have a document with hundreds of tables with more or less the same kinds of content, figuring out the difference between the two. Again, very hard to, that's the task specific model.
You Have a table specific model that extracts the information from these tables and provides queries and understanding and all that stuff. Yes. Most corporations have gazillion tables, uh, spreadsheets, let's call it.
That's Exactly why we had to go build one, because as, um, non cool as it sounds, oh God, yeah. That problem is a problem that every enterprise has over and over and over again. And you need to be precise.
I can't miss three digits and say, oh, I got the four digits. Right. Whoa.
And if they, And if they're operating a sensitive environment, this could be on-prem gonna be air gap. Yeah, Completely. Yes.
So we, so to that question, we are deployed in the cloud, we can deploy in a customer's cloud and we can also deploy into completely our gap environments, right? In fact, we've been in production in an air gap semiconductor manufacturing facility now for over a year. And we do, it's a, an augmented technician assistant that's running next to Absolutely.
So really quickly, um, the other thing I want to mention is we do not look at domain specific models just from regular benchmarks. And um, I know some of you are from the benchmarking world, benchmarks are great, uh, but they can also be weaponized to be used against somebody else. We Don't allow that, right?
So, um, and looking at how do models perform from multiple angles is critical to us deploying into production, right? And so what this plot is telling you is this is the energy model again, across 10 different dimensions where the tests actually came from the experts and the model is being tested across all of those dimensions. The closer it is to a hundred, of course it's answering many questions correctly.
It's not that the blue that is the domain specific model beats the general purpose model in one or two or three things. It's across the board significantly better than every other model that's out there. 1 was the state of the art.
We actually launched it in GTC in November, in uh, March of this year. 3 K, then LAMA 4K, then G-P-T-O-S-S came, then GPT five came. You can see the benchmark across all of them.
Every single general purpose model actually significantly improved the capability of the previous general purpose model capability. In spite of all of that improving, you can see that the domain specific model held its own. We didn't retrain it.
It's the same model from us, right? And that's important because for us it's not just about building a mote, it's about building resilience. Because you can go build general purpose models, which will continue to improve.
We will benefit from that because we learn from the architecture that comes out. But then if you're not going after the tasks that are specific to that domain, you're not going to actually beat that model. That's important.
Right? And we have similar plots like that across every other domain that we have, whether it's finance or supply chain or automotive or, uh, semiconductors or design or any of those models as well. Okay.
How would you say task specific? You mentioned spreadsheets, but uh, aerodynamics to some extent is a task that multiple verticals could potentially use. Yes.
So those are the sorts of models. How many task specific models do you have? So The, so the aerodynamics would for us would be a domain specific model inside aerodynamics.
If there is a particular task, for example, in that particular field, if you have to understand airflow diagrams, okay, that would be a task. We might either build a particular model for that or that might actually be a subpart of the, the domain specific. So the customer process onboarding is, so let's say I'm a new vertical.
Let's say I'm a bakery and I don't want to have a, a domain specific models for bakery. Do I pay, uh, professional services for you to do this? Or Typically no.
So if we walk into a vertical and we have a customer in mind, or there's a customer who says, if you had this model, we would actually pay for it. We would actually do the investment to build the model. That's because we want to make sure the ips our, if we build any model with the customer's data set, that model is the customer's model.
We typically don't do that. What also happens is customers have come in, for example, cybersecurity. We had a system analysis model that we had.
We also built a cybersecurity model that understands cybersecurity from the, the domain specific standpoint. The customer came in and said they have cybersecurity needs. They want us to build a model for them.
So we use our cybersecurity model, we use their data set to augment that model. But the final resulting model was theirs, right? So we very rarely do that.
The only reason we did that for that, that customer was it was a, a very large cloud provider and it was a validation of our technology that we could do. And they would then subscribe to the product. Because remember the premise is a model is only a means to an end.
We need multiple models working together. You need an intelligence layer to actually get to the outcome.