Affordable on premises LLM training and inference with Phison aiDAPTIV+
Phison’s show their aiDAPTIV+ technology, designed to make on-premises AI processing more affordable and accessible, particularly for small to medium-sized businesses, governments, and universities. The core of their innovation lies in leveraging flash storage to offload the memory demands of large language models (LLMs) from the GPU. This approach addresses the growing challenge of limited GPU memory capacity, which often necessitates buying more GPUs than needed, primarily for the memory capacity.
Phison’s solution enables the loading of LLMs onto high-capacity, cost-effective flash memory, allowing the GPU to access the necessary data in slices for processing. This significantly reduces the cost compared to traditional deployments that rely solely on GPU memory. The company is partnering with OEMs to integrate their technology into various platforms, including desktops, laptops, and even IoT devices, with a focus on providing pre-tested solutions for a seamless user experience. They are also expanding their partnerships to include storage systems.
Beyond hardware, Phison also addresses the knowledge gap in LLM training by offering educational programs and working with universities to provide students with access to affordable AI infrastructure. Their teaching PC, offered in partnership with Newegg, aims to democratize LLM training by making it accessible in classrooms. The company’s efforts focus on fine-tuning pre-existing foundational models with domain-specific data, allowing businesses and institutions to tailor AI to their unique needs and keep their data private.
Presented by Brian Cox, Product Marketing Director, Phison. Recorded live in Santa Clara, California, on April 24, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/phison-technology-presents-at-ai-infrastructure-field-day-2/ or https://techfieldday.com/event/aiifd2/ for more information.
Transcript
I'm Brian Cox with Fon. Uh, we are, uh, headquartered in Taiwan with US headquarters here in Silicon Valley in San Jose. And we're gonna be talking about an application of our flash memory technology that actually unleashes the economics and the scalability for AI processing, particularly in LLM training, as well as inference.
A little bit of background further on f we've actually been around a long time, 24 years, uh, over 4,000 employees. Most of those are engineers in Taiwan. Uh, in fact, we helped pioneer the first, uh, uh, USB thumb drives and other flash memory technology, and are a key supplier of many of the, uh, key components that go into a lot of the SSDs in the industry.
In fact, we are the largest supplier of the controllers, as an example, that go into all the SSDs that you're using, uh, both in your personal computers as well as in data center equipment. So we do this in a variety of different, uh, applications. And we have a large customer base, which I'll give a little glimpse of here.
You'll see Fon technology in many of the systems that you see on the left, uh, name brands that you recognize and it's growing. Uh, each day as we get more and more design wins, but we also have technology exchanges with a number of the other players in storage that you see in the center. Uh, we're buying Nan Flash from them.
They're buying controllers from us and firmware and such. So we're very deeply ingrained with the whole ecosystem in regards to flash memory, given that we have control of the firmware, control of the controllers themselves that manage the reads and writes, and the endurance and the security of the drives, we're able to design for very specific use cases as well. So for automotive, we supply flash memory and storage for that.
For other industrial uses. We also supply the storage that NASA uses. So the lunar lander, the mission to Mars, that's Fon technology.
So we were looking at other ways that we could help across industry with our knowledge of flash. And we realize there's a big challenge that we have within the adoption of ai, and that is in regards to the capacity of the memory to handle these ever-growing large models. And we can apply flash as an alternative to simply being constrained by the memory that's on the GPU card.
We can now unleash tenfold more capacity at very affordable prices. We go ahead and give a little bit of the history of how we got into this. Fon, as a company wanted to adopt ai, our engineer said it would save us a lot of labor time.
It would give us, you know, better insights about our processes, and we're loading all our documents in there. But what we found, and this is private, this is our designs. We didn't want to put these out on the public, uh, AI platforms.
We needed to have that in-house, but to do that mean meant that we had to buy our own infrastructure to house this. And so when the engineers said, well, if we're gonna do this in house, you know, here's the equipment that we need, and the price tag came up to well over a million dollars, and our CEO goes, that wasn't in the budget. Uh, can you guys go back and figure out a more cost effective way for us to apply AI within vison and serendipitously, which happens a lot of times here in Silicon Valley.
Um, your, your knowledge meets the opportunity given that we are in the memory industry. And that was the key driver in regards to the cost of these systems. We figured that instead of applying or loading the model solely into the VAM, the high speed memory that's on the GPU cards, and having to have many, many, many of these cards to get a big enough memory pool, we could instead load that onto flash, which has very high capacity and is very affordable.
So with some smart engineering in the middle, middleware handled all the memory management. So we could then provide slices of that model to the GPU as it's needed for processing, but being able to load larger and larger models because of the flash capacity. So it completely has changed the economics of GPU processing.
Now you can load the largest models and scale that easily as the industry continues to move forward and be able to do that and afford it with flash. So now it's no longer just a domain of Fortune 500 companies, but small medium businesses, uh, state and local government, uh, universities can now afford to have their own AI infrastructure to run their own models, to apply their research. And it now democratizes AI for all of society.
So you can see there's a number of different, uh, applications that people are looking at across industry. This is just one glimpse. You know, we were looking at it from product creation, also manufacturing, but everything from, uh, financials with market evaluation, uh, sales prediction, all types of different use cases that many of you have heard of, we now wanna make that so businesses of all sizes, uh, can afford that.
But one of the things that we've run into, there's a few constraints. We talked about just the size of the models, and we'll come back to that. But also we found that there is not many people that understand how to train these models, these large language models or LLMs, uh, that is rapidly evolving.
People, you know, don't have a good comprehensive introduction to how to do this. And then how to build their own models. So one of the ways that we have determined is that if we can provide infrastructure that's affordable to universities, they can give their own platform access to all of the different students.
Unless you're a very large university where you have, you know, endless budget and you can have enough infrastructure that every student can get access to the systems that you know, have the LLMs, unless you're one of the very largest universities, you're gonna actually have to have all of your professors, all of your researchers, all of your students share time on that. And it may be days, weeks before you get your slot on that server to be able to do the training. If we can bring the cost down for the large language model training and research.
And so you could have it on your desktop, you could have it in the form of a laptop, you could then train models everywhere in the classroom, in your dorm room, in an office, anywhere you would like. So now everybody has access anytime they need it in order to learn how to do AI processing. So we saw that this was a big constraint.
We need to enable the universities to be able to, uh, provide the students the access. And this gives a picture here in this chart of the growth of RAM that's used to house those models. And that's growing about, it doubles about every two years, but the rate of growth of the size of the AI models, the large language models, is going up over 400 times every two years.
And we were talking about large models just a couple years ago of 70 billion parameters. Then it went to four oh 5 billion parameters deep seek few months ago, 671 billion parameters. And now meta is talking about a 2 trillion size model they call bmi.
So, and that's anticipated roughly in about a year. We need to have infrastructure that scales to those large sizes and also be able to afford it. So the technical challenges that we see, we talked about this, the knowledge base we need to enable universities to be able to train the next generation of AI engineers, but also there's some technical challenges we talked about in fine tuning the rapid growth of the size of the models, uh, there's just not enough memory.
It's, it's difficult to scale, it's expensive to scale. And then on inference, which is where you actually get for the everyday user, you get the value out of it. When you go to chat GPT and you enter your prompt, right?
You're getting information back and then you say, oh, I want to learn more about this or put it in this format, and you do all these sub uh, subsequent prompts. We wanna make sure that that's a good experience. And what we've found is that with limited memory, when you start doing those prompts and you do successive prompts off your initial query, you run out of memory bad user experience.
So what we found is by providing the added capacity with flash, we now enrich that experience because we don't run out of memory. So we're giving value to the greater population of people who are using these models inside a company. This is really meant for on-premise, but also for the people who are building the models with the fine tuning, the added memory solves both challenges.
So I'm gonna give a little bit of a product overview. We do have some of the devices that you see, uh, up there on the screen right here in front of us. So these come in different form factors to fit everything from IOT devices to laptops to workstations to servers.
And behind the delegates, we actually have a display table with a number of these products provided by our partners. So we can fit this flash memory for AI purposes in really every form factor. A little bit about the, the technology and my colleague Sebastian will be coming up in just a couple minutes and he'll go a deeper dive into this.
But basically what we're doing is, I explained earlier is that you have the high speed memory, uh, goes under a couple different names. H-B-M-G-D-D-R generally called VAM, that's on the GPU card. In fact, if you look at the GPU card, in many cases the most expensive component on the GPU card is not the GPU, it's the memory.
And I was just at NVIDIA's GTC trade show, and I had customers come up to me and say, you know, I was needing to load large language models, you know, on premise at where I work, and I had to buy multiple GPU cards. And he goes, I didn't need the GPUs, I just needed the extra memory for a big enough memory pool. And I said, well, that's a pretty inefficient way to do it.
And he said, yes. He goes, I wish I would've known about your technology earlier because he could then buy the number of GPUs he needed and then get the flash to give him the memory capacity. So what we do is that instead of having the whole memory pool over on the, solely on the GPU, is that we bring in flash as a cache, load the models there, and then take slices of that to feed the GPU as it's needed for processing.
And we can do this now at one 10th the cost of a traditional deployment that you see on the left. And this gives a, an example of the footprint saving. So if I am using fewer GPU cards, which are represented by the rectangles there on the left, and a little bit on the right, we shrink the number of GPU cards that you actually need because now we can load the models onto flash.
And this saves you upfront obviously capital cost purchase, you need fewer systems, fewer GPU cards, and it also saves on some operational costs. So fewer systems, there's lower maintenance costs, lower support contracts, less footprint. You save on that as well.
So Brian, yes. Um, return group that interesting about the GPU and okay, so your premise here for the statement here is that I'm buying GPUs because I need the memory. Mm-hmm.
In what cases is that not true? Uh, is that, I mean, are, are, I'm, and we have, we have maybe somebody around the table knows this because they're smarter than me on the, the performance stuff, but what I'm looking at is say, okay, so we've been, we constantly talk about how we need to feed, feed the GPUs and keep the GPUs busy. Mm-hmm.
Um, and um, for all that reason, so I'm like trying to say, okay, so am I not framing my, my title? You're doing fine. I'm, I'm stumbling all over myself right now.
Um, I would think, is it a corner case that we're having to do that, that this has happened to happen or what This is happening? Well, we're, we're running out of memory footprint to be able to handle the growing size models is happening more and more and more. And the economics of it are happening every day.
And I understand the memory pool. I mean, we're trying to bring this memory pool, but how much, how often is it, or how significant is it that people are buying GPUs because they don't have enough memory? Uh, that seems like that asks backwards That it does.
And it's happening very often. That's the challenge and that's the opportunity that Fon saw that we could help alleviate that problem. Now, if you have a small model and it fits within the GPU memory, that's fine, but you'll see some chart, you know, in a few minutes that show that where you fit it within the memory pool that's on the GPU card, you're fine.
But then as you grow that model, you then run out the memory and it, you can't even do, do the training. It doesn't work at all. Okay.
So let me phrase, have you guys looked and looked at the market and established what percentage of the deployments that are doing AI training or inferencing are hitting the wall in this manner and therefore buying more GPUs? I, I don't have precise numbers. Okay.
To be honest. And I figured you, because, you know, it'd be hard to do that. Yeah.
I was watching this, uh, uh, webcast by a number of venture capitalists and they were making their predictions what are gonna be the most hot topics in 2025. Mm-hmm. And they're going around and some of the typical ones that you hear, and one of the venture capitalists and said the highlight of what's gonna happen in 2025 is the lack of memory capacity for AI processing.
And absolutely agree with that. Yeah, absolutely agree with that. But I, I don't have a specific number.
I don't know that anything, any research has been particularly published on that yet. Okay. Uh, but it's a widespread known issue.
Okay. Uh, so we'll, we can talk further as we go through the presentation. Cool.
Thank you. But no, I, I wanna make this interactive. So I know I've been talking quite a bit and Kimberly, thank you for interjecting with a question because I want to hear what your perspectives and, you know, your experiences are.
So I'll, excuse me, I'll ask the follow on Jack Pauler paradigm. Yeah. Um, Kimberly alluded to it, but is this solely a training problem or is this also current inference?
Uh, it's also an issue for inference. Totally both. Yes.
So, and we'll talk a little bit about that inference performance in just a few minutes. Okay. 'cause that opens up the market significantly.
Oh yes, it does. In fact, we started on this path initially identifying the fine tuning problem. Our customers came back to us and said, we have an inference problem too.
We ran the models, and particularly for smaller devices that don't have a big memory footprint, by adding more capacity Mm-hmm. We dramatically increase the inference, uh, performance and experience. And we will, we'll show some charts on that.
Right. Then I'll ask the, the other follow-ons, which is stupid or not stupid question is, is this a problem that could also be solved by redesigning a GPU card? Is it a more memory footprint on the physical board or is it the GPU?
Can't physically address any more memory. Uh, and I'll have, uh, my colleague Sebastian talk about that. But I believe that the constraint about the amount of memory that's on the GPU card, I think is simply the, uh, the economics and the physics of the space to be able to fit all that memory alongside that GPU card.
And you know, why it's limited. The, once again, you got that earlier chart, the density of memory is not growing as fast as the growth of the size of the models. Yeah.
So, I mean, memory density is growing, but not as rapidly. Okay. And you'd end up having probably pretty massive GPU cards, which then creates all kinds of thermal problems, power problems as well.
So if, that's a great question and maybe Sebastian can talk further about that, uh, when he comes join me in a moment. Uh, what other questions? Okay, I'll keep going and feel free to just, uh, wave at me and we will answer any other questions that come to mind.
So this gives a perspective on the breadth of not only platforms, but also different price points. So we were showing at the NVIDIA conference an IOT device that was using, you know, this very small form factor looks like a stick of gum, uh, uh, SSD. And so with IOT devices, this can go in, you know, a myriad of different locations.
Everything from digital signage to your, uh, automated vending machines to all kinds of sensors on a manufacturing line. We had a lot of interest from people coming by our booth at GTC wanting to deploy in an iot edge kind of environment. So how much Storage is on that?
Hi, Denise. You, Uh, this could come, it is currently available in, uh, 320 gigabytes all the way up to two terabytes. And we have on our roadmap to soon be going to eight terabytes on this.
Pretty crazy in regards to the size that you can get with this. And then we have other form factors as well. Yeah, that's very cool.
Uh, so these are the platforms that we've already validated our technology in. And, uh, here's a number of the vendors that we have their products on display. They, we either bought or we had, they loaned us their gear.
So you can get this, you know, today, uh, you know, in from New Egg, for example, you can get a desktop pc or you can get a high powered engineering workstation with our technology integrated in, you can order it online right now. Okay. Hey, Brian?
Yeah. Uh, Brian Martin, signal 65. Um, when you say training here, are we talking training base models?
Are we talking fine tuning or additional training? What do, what do we mean by training? That's a great question.
So once again, I, I should, uh, make sure everybody's level set. This is really for on-premise. Mm-hmm.
Use, because you may have very, a variety of reasons why you don't wanna put your models in the cloud. Lots could be privacy, it could be for cost predictability. A lot of reasons you may not want to have this up in the cloud.
So, uh, what, and so your question again, sorry, was When you, when you say training on Oh Yeah. The training. Okay.
Foundation model training. So the, or fine tuning an existing foundation. Yeah, it's, it's fine tuning.
So the, you get the, um, foundational models. You either get the ones that are commercial ones, the chat GPTs, the, you know, quad from perplexity. There's a lot of them that are out there.
And then you get the open source versions like meta Right. Offers an open source foundational model called wma. Mm-hmm.
You can download WMA free of charge, you can download Deep Seek, which is also free of open source, free of charge, but it only has general knowledge. Mm-hmm. So I'm old enough, I remember the old encyclopedias, right.
I'd go to the library. There's this record of the knowledge of humankind sitting there in multiple volumes on a bookshelf. Right.
However, it knows nothing about me. Mm-hmm. It knows nothing about my business and or even in specifics about my industry.
It may not even have in there. So that limits the value that large language models can provide to your business. It doesn't give you an edge, but if you can bring those models, download lawn mow, or download deep speak mm-hmm.
And then augment them or think I'll train them domain specific with additional domain specific information about your business. Could be product plans, customer lists, contracts, all kinds of different things that's unique to you that's private. You don't want up in the public cloud.
But that would help your employees get better answers back because now they have, we've added to that encyclopedia the knowledge specifically about your firm, your industry. And it makes it that much more powerful. So that's, there's a number of techniques that are used to go ahead and augment that information.
Uh, retrieval, augmented, uh, generation rag, you'll hear about full model fine tuning. You'll hear about QA fine tuning. It's a very, you know, dynamic time in the industry as they're, you know, looking at the best ways to combine these to get you accurate results based on all the data that's been fed to it.
So just going back to the small storage form factor that you had there. Yeah. How actually does it work out at the edge?
So where would the storage be? So if you've got sensors, um, where would that would plug back into a field server or what would did? Yeah.
So I'm gonna ask my colleague Loretta, 'cause she's standing right next to a device. If you could grab, there's a device that is very popular out there from Nvidia called uh, the Jetson family of IOT devices. Mm-hmm.
Um, and there's other versions as well of small form factor computers. So this is an example of a device that's also gonna be, it's available for many different manufacturers where they wrap an enclosure around it. And so it can go into these remote locations.
It could be on a manufacturing line, it could be some remote sensor on an, in an oil field. This does the processing for that. And so you can see where we've actually installed our memory that's used as a cache to hold the different models to be able to do the AI processing.
And so those would be ruggedized and they would be within. Correct. Would they sit within a larger server or how would they then connect back to then?
Do they go to interim, uh, server or router and then back to the data center? Yeah. Uh, yeah, it would, you know, be connected over the internet, but also it works where you may not have good internet connectivity and you could do your processing locally.
Okay. That's an, that's a really key advantage if we're in some far Remote rather having send it all the way back. Yeah.
Part of, I don't know, South Dakota and made much internet. This is the way you could do the processing out there locally. We're gonna get letters from South Dakota now.
Say What? Well, wide open spaces. Beautiful.
Yes. Yeah, there you go. Okay.
Brian, I do have a question. Uh, Matt from Osmium Data Group. There's an interesting thing here in this slide.
I mean, do you have servers on the right side? But I see laptops, desk desktops, workstations, and even IOT devices and it kind of makes me think, you know, what is your intent in there? If I ask it very bluntly, I mean, is your intent to kind of partner with OEMs, you know, like laptop manufacturers, the stuff manufacturers and so on?
You kind of have your technology directly embedded into the devices? Or is it something where you see that customers, I don't know, companies, individuals will go and get that themselves embedded into the product somehow, like buying one of your SSDs or whatever you do. Mm-hmm.
So our initial go-to-market motion is to work with system vendors that they will integrate our technology in there. 'cause it's more than just, uh, the SSD, we'll talk more in detail about this, but there's also the middleware that does the memory management and you could interact with that at a command line or we also provide an all-in-one tool set in a graphical user interface that goes with it as well. What we do with these manufacturers is that they pre-install this on their equipment.
They test it. So when you receive the shipment and you unbox it, it all works. Uh, the challenge, if you just provide upgrade kits out there, you don't know what the configurations are and we want, you know, so we want to move carefully before we just make it available, you know, in upgrade kits because they may put it in some custom built PC that the motherboard is incompatible with the dry with something a lot of unknowns.
But when we have it shipped out from New Egg in a workstation or a desktop pc, they've already pretested it. Main gear, a high performance laptop manufacturer will be offering that this quarter. They've already integrated it, it just runs.
We're working with a number of iot, uh, vendors that will incorporate the Jetson device in their systems. They've already tested it. We were gonna have a server here actually from MSI, but it's not even arriving until later today.
So I couldn't bring it with us, but I would've had that on display as well. So they were all integrating testing, making sure there's a wonderful customer experience. Maybe down the road we may offer upgrade kits, but that's not the plan right now.
So I did also want to, uh, give a preview of an expansion of a strategic partnership with yet another platform category. This is in storage systems. So we have a partnership with many different storage vendors, of course offering our standard SSDs for, you know, regular storage.
We're now going to be working with, so one to offer AI processing as part of their storage system deployment. So this is like an expansion of the partnership, but the actual details of the product, uh, when we will come out with a formal announcement about that specifically later this quarter. So this is once again adding to that previous slide beyond just servers in the data center.
Now also storage systems in the data center will have AI processing integrated as part of their, uh, capabilities. Can you talk more about what you're doing with them? We will, when we have the detailed Okay.
Product announcement later this quarter. Okay. But that's coming.
Okay. So, uh, I talked about one of the constraints earlier. Besides the technology, you know, the constraints on being able to scale the memory, the cost of doing so.
Another constraint I did mention is just a lack of knowledge. There's not enough people out there who know how to confidently train large language models. So we are, have worked with Newag to come up with a, what we call a teaching pc, can imagine this available for every student in a classroom where they have access, all of them to their own individual systems and they can do the training right there.
You now bring the price point down where it's affordable. Instead of spending millions of dollars, you spend a few thousand dollars per student outfitting a classroom, and now they can do the large language model training and the inference, uh, research right there as part of their classroom education. So we see that, uh, this goes, there's been a I PCs out there, but all they did was just do inference.
Well now we go beyond just using an LLM for inference, but now learning how to train in LM with your own infrastructure. In addition, we realized we needed to take a step forward in addition to universities offer our own classes. So we've already been doing this in Taiwan.
We've run a number of classes, had a number of students already gone through it. Obviously that's, there's a course description there on the right in Chinese. Uh, we also have, uh, a joint venture with a company called My Storage in Malaysia.
And we're offering courses there. We're gonna be bringing this also to the us So it's more than just theory. It's more than just learning what is a ai, what is an LLM, what is inference.
It's more than just that. It's actually hands-on workshops where they have the opportunity to do it on their own, access to their own machines that are right there as part of the classroom training. So this is gonna help accelerate and alleviate that other bottleneck, which is just lack of knowledge within the industry, how to do AI processing processing.
Quick question, are you referencing when, when we're talking about training LLMs, are you referencing building additional knowledge into foundational models or fine tuning foundational Models? Fine tuning. Okay.
So there's a lot of research, a lot of money by very big companies from, you know, Google to meta to Elon Musk Right. Doing the foundational models. But once again, there was, are just general knowledge, but we want to do is now make it affordable and accessible by individual companies, state and local governments, universities, that they could do their own research and augment those general knowledge models with unique information that they may wanna keep private and, but make the LLMs and AI much more useful.
Mm-hmm.