Our Major Takeaways from AI Infrastructure Field Day 2 – Tech Field Day Takeaways
In this episode of Tech Field Day Takeaways from AI Infrastructure Field Day, Alastair Cooke, shares key insights from the event, highlighting how AI infrastructure must align with each phase of the AI pipeline—from data ingestion to training, fine-tuning, and inference. Training phases demand high GPU and memory throughput, while inference, often the dominant workload, benefits from strategies like SSD caching to handle growing model sizes. To ensure performance and scalability, network design must prioritize high throughput and traffic segregation between GPUs, storage, and applications.
Alastair Cooke: https://techfieldday.com/people/alastair-cooke/
Event Information: https://techfieldday.com/event/aiifd2/
Follow Tech Field Day:
Website: https://techfieldday.com/
LinkedIn: https://www.linkedin.com/company/tech-field-day/
X/Twitter: https://x.com/TechFieldDay
Bluesky: https://bsky.app/profile/techfieldday.com
Transcript
I'm Alistair Cook. I'm an event lead here at Tick Field Day. And these are my takeaways from AI Infrastructure Field Day two, AI Infrastructure Field.
Day two was a massive event. It is the largest tech field day event I've ever been involved in. We had a full four days and a few themes along the way.
One is the, the sheer complexity of building AI solutions and building AI into your applications. And we saw that through the suite of products and developments from, uh, Google as well as through most of the other presenting companies showed us assistance on building these complex solutions, building them out to fit your requirements and your requirements at different stages in the lifecycle of your AI infrastructure. Network design for AI is complex and we've seen this right from the beginning of running AI events at Tech Field Day.
A lot of the time it's the network design to get maximum throughput and particularly to keep feeding the GPUs and the GPUs are really the most expensive part of your AI infrastructure. So the network design is about making sure you are getting the data into those GPUs and we see some pretty complex and high performance designs around segregating out that GPU to GPU traffic from storage, traffic traffic, and from your application request traffic coming in the front. So network design is absolutely crucial as you're building out an AI infrastructure.
AI model sizes have been getting bigger and bigger, and this leads to some significant challenges, particularly if you can't continuously replace your GPUs with larger GPUs that have more memory attached to the more high bandwidth memory. One of the themes we saw a little through our, uh, sessions at AI infrastructure fields, I was finding ways to optimize this. One of the themes was using SSD as a case to hold some of the model.
So you don't need to use A GPU that can hold the entire model in memory. It's important to recognize it's not just the model, but it's the entire context and all of the session information around the actual use of the application when you're going through this inference stage. I don't think we're seeing SSDK as being a thing during the training phase, but definitely during the inference phase, which is probably going to be the majority of the actual workload for most AI applications.
AI as a pipeline as a series of stages was a big theme for us, seeing that there are different types of workload at different phases through that, uh, that pipeline. So seeing the data ingestion phase where we're bringing data in to start training and then the training stages, these are a very foundational beginning of the process stages and they have very specific requirements around data transfer and processing. So data ingestion and processing is a bit of networking coming in, but quite a lot of compute power being applied.
And then during the model training phase and equally the fine tuning phase, we're feeding huge amounts of memory, uh, with huge amounts of data in huge numbers or large numbers of GPUs. Now all of these phases are really just the preparation for doing the actual work because the actual work is doing inference. It's building the AI into our application and using often retrieval, augmented generation or, uh, the fine tune models to actually enrich our application and deliver value.
And that inference phase is again, quite different in its requirements. It does again require CP, uh, C-P-U-G-P-U, but the GPU requirement is a little lighter than the GPU requirement for that training. And fine tuning, recognizing that as you're building a solution for ai, building an infrastructure for ai, you need to build slightly different infrastructure for these different phases.
Now hopefully you've got a fairly general purpose infrastructure and it suits those different phases, but definitely being aware of the different phases and maybe you are not actually doing all of them. Maybe you are not doing the initial training and you're starting with a foundation model and just doing fine tuning or retrieval. Augmented generation AI infrastructure is definitely a challenging area for a lot of organizations.
And this is reflects the fact that we had a lot of different approaches to building AI infrastructure at our infrastructure field day two, and that we really did see a lot of real world application, a lot of maturity in some of these, these applications and some views to running your AI inference on premises rather than maybe running in cloud as we thought all AI was gonna be done earlier on. We were delighted to welcome Fon for their first presentation showing both their AI adaptive technologies, but also then new pisca range of SSDs that you can buy directly from new. I was also really happy to have Nutanix back Nutanix presented at my very first Tick Field Day event, uh, back at Tick Field day nine, and this is the first time they've been back with us.
So welcome back Nutanix at AI Infrastructure Field Day two. I was delighted to have really strong support from the wider RUM group for this event. Uh, Brian Martin from uh, the Signal six five has uh, joined and this was his first event as a delegate and he got to experience the fire hose of learning that is Tech Field Day.
Uh, Kaley Bates, uh, was, uh, happy to join us as well and I really enjoyed spending some time with Kaley, also with the rest of my delegates. Uh, guy K and Mitch Ashley also joined us for our day at Google. So thanks to the wider Futureum group, team.
Team. It's been really good to have you with us. I'm Alistair Cook and you can find me on LinkedIn as well, Alistair Cook.
Just search for me. You'll find me. Uh, you also can find me on, uh, many of the social medias as Demi Tass nz.
You'll see lots of video from AI Infrastructure Field Day two published on the Tech Field Day YouTube channel. And of course, you can find more information about all of the presenting companies and all of the delegates on the Tech Field Day website. Thanks for watching this episode of the Tech Field Day takeaways series on the Tech Field Day plus YouTube channel.
If you enjoyed it, please subscribe, like, share the video, maybe, uh, add a commentary on your thoughts on the AI infrastructure field day, uh, in the comments below, follow Tech Field Day on X, Twitter, blue sky and master it on For more updates. You can watch all the presentations for the videos on the tech field, their website, and on YouTube. Thanks for watching and we'll see you next time.