The Language Model of Tomorrow: Charting the Course of Generative AI with Richard Shan at Techstrong Con 2024
In this presentation, we explore the progression and trends of generative AI, focused on large language models (LLMs) and small language models (SLMs). Beginning with an overview of their characteristics, use cases and values, the talk then forecasts key advancements in the evolution of LLMs, the rise of SLMs and their synergistic interplay. This exploration extends to examining the societal and economic impacts of these movements, addressing both the potential and the technology challenges they present. Crucially, the deep dive underscores the importance of navigating the associated risks, emphasizing the need for interdisciplinary collaboration, responsibility guidelines, and explainability considerations in generative AI development. By advocating for sustainable practices and effective governance, the analysis highlights the necessity of balancing technological innovation with ethical responsibility to ensure these advancements benefit society as a whole.
Transcript
Hi everyone. Good morning, good afternoon, good evening. From wherever you're joining us from.
My name is Richard Shan and in today's presentation I'll be discussing the evolution of generative AI and specifically the field of language models. So let's start off with a quick overview of today's presentation. I'll start off with an introduction to today's topic.
We'll talk about some of the current trends in the general generative artificial intelligence or gen AI space. Then we'll take a look at language models. I'll give you a quick overview and some definitions and then we'll talk about large language models versus smaller and small language models.
Then I'll start drilling down into small language models specifically I'll talk about some of their key appear key appealing characteristics in terms of performances and applications. And I'll explore some examples and use cases. We'll then talk about how small language models can hold up to these larger and bigger language models.
We'll then discuss how we can look forward into next steps so that we can increase the use of small language models and a couple of challenges that we have to address in the generative AI space. And finally, I'll close out what the call to action and talk about some important next steps for the field at large. So let's start off with the introduction.
Generative AI is a really hot and grow and growing industry. The industry's value is projected to triple by nearly 2030 and at least a third of the companies in the industry use generative AI in 2023. Outside of the workspace, Chachi BD has transformed billions of lives in how people interact with technology across the world.
So let's talk about generative AI. At its core, generative AI refers to a bunch of algorithms that can generate new content from text to AI systems. These systems are just analyzing data that was already there.
Rather generative AI systems are new and that they're able to create data that did not exist previously. Whether it's in healthcare, finance, entertainment or education, gen AI is revolutionary revolutionizing how these industries can operate. It's having a bunch of abilities to generate novel solutions and its ideas are transforming business models and customer experiences.
Under the current umbrella of gen AI technologies such as like machine learning, neural networks and NLP or natural language processing are playing pivotal roles in their expanding and seeing a lot of use, they form the backbones of systems that can emulate these humanlike creativity and reasoning. In gen AI systems, gen AI potential extends, just extends beyond what it's currently being used for. It promises advancements in things like medicine, telehealth and personalized treatments.
It's able to automate content generation and create intelligent virtual assistance and more. But let's address language models in particular. Now, language models and gen AI aren't necessarily the same thing.
A language model in and of itself is basically a tech a text prediction model similar to if you're ever typing on an iPad or a phone and there's a little bar at the top that says, oh, this might be your next word. That's essentially what a language model's doing. When you query a language model and ask it something like, why is the sky blue?
It essentially faces its response on its train data. It'll have a bunch of websites and resources that it already has been trained upon on knowing why the sky is blue and then it'll generate one word at a time and it's kind of like filling in the blank. It'll generate one word at a time after its previous word and it'll cohesively create a response based on its training data.
To answer your question of why the sky is blue, language models are just a singular subset of generative AI and inside language models we still have more subsets, specifically large language models and small language models. Large language models, as their name suggests, are inherently large. They have more training data and more powerful processing and they're usually for broader applications and can be applied across the spectrum and can answer topics, uh, and can answer questions on a lot of topics.
For example, if I were to query chat GBT, it could tell me why the sky's blue, but it could also tell me how a single cell divides and creates more copies of itself. However, small language models are fundamentally built with the goal of keeping the size as small as possible. That's usually because these smaller language models are built to run on a small device locally and they don't wanna use much processing power or computational power.
These smaller language models are usually adapted for a specific area. For example, a small language model might not be able to to answer a question about rocket science, but it might be able to answer a question about, for example, 18th century English or French because it only really needs to be in a small and domain specific area. That means these smaller language models are able to be a lot smaller in terms of training set size in terms of computational resource consumption.
These smaller models really only need to know a lot in its specific field and thus becomes not only a lot less resource intensive but also a lot more accurate in its specific fields. When you compare them to sweeping and broad large language models, their smaller size and their domain specificity is what really makes these smaller language models appealing as an alternative to just applying a larger language model to everything across all applications. So let's look at a comparison.
If we look at size alone, as the name suggests, large language models are a lot, lot bigger than these small ones. 8 trillion parameters large. However, if we look to a smaller language model being developed by Microsoft FI two, that's just under 3 billion parameters.
In fact, it only has put 7 billion. This small language model is over 500 times smaller and a lot easier to train and process compared to a larger language model like the GBT series. Then on training sites where the pattern repeats itself, again, the GBT series is trained from data called all across the web and it only has data as recent as April, 2023.
However it is from all across the internet and all across a bunch of resources wherever they can get data from versus, although it's, although this application is not specific to FI two necessarily, many custom built small language models are only domain specific. That means that they only need to be trained on a lot of data in a specific area of expertise. For example, if I, if I deployed a small language model in a retail store, I might only need to train that language model in particular on say how how many refrigerators have been sold in the past or how consumers like their refrigerators, their features, and uh, past statistics on selling refrigerators.
But that means is that you have a lot, you have a lot smaller of a training size 'cause it's specific to one topic. However, you also get a lot more accuracy in that one topic because you're able to hyperfocus your language model on that one topic. And next is on resource requirements.
To train a large language model is incredibly resource intensive, not only to query but just to train one, you need thousands of GPUs and a GPU farm, lots of electricity and and lots of electricity and lots of computational power overall. And to keep one running after you trained it, it needs constant maintenance and you need to be watching over every single query. However, small language models offer an advantage in this area.
Smaller language models are usually designed to be run on the edge or on a really small device like a phone or some sort of microcontroller. They don't need a lot of processing power, whether you're trying to query it or whether you're trying to train a small language model. It's usually really easy to host locally.
They're also really low maintenance overall 'cause you don't need to monitor thousands, thousands of different variables to see which, to see what happens and if anything went wrong. And lastly is speed. I think if you use chat GBT, we've all experienced a slowdown at some point, whether our WiFi's been pretty bad or we ask chat GBTA complex question that doesn't really know the answer for chat, GBT will slow down in generation as it grasps through a straw trying to find an answer.
Even if you have a really bad wifi connection, it can take a couple of minutes just for it to answer one query instead. Small language models have an advantage in this area too because they only need to have knowledge in one specific area. Small language models never really have to jump through hoops to reach a complex answer.
Oftentimes because small language models are specifically designed to be easily run locally, they're often run locally at the edge, which means that even if you have a really bad wifi connection, it doesn't matter. You don't need to rely on wifi latency to make sure that your language model is able to respond in a fast and timely manner since everything is running offline. So we compared some differences between small and large language models.
Let's drill down into some of the use cases and specific characteristics of small language models. Small language models have a lot of uses in a lot of applications due to the fact that they're able to be run offline a lot smaller and able to be put on a variety of platforms. For example, we can put small English model on a bunch of chat bots because these small language models are real time and personalized, which is perfect for the application we're trying to use.
Oftentimes these smaller language models are also a lot less power intensive and it can drain the battery a lot slower than using a large language model as well. Because these small language models are hosted locally, it means that we can put them on a on device chat bot without it needing to necessarily be connected to the internet. For example, if we put a small language model on a smart watch, this would be great because it means we would have a low power usage for a really long battery life and we could still run our language model without internet, which is not always guaranteed.
But secondly, small language models are also great in scenarios where we need real time or live decision making. For example, if we're dealing with the stock market and a couple of thousand and a couple of milliseconds can mean the difference between a massive loss or a win. Small language models have really fast speeds because of their domain specificity.
With small language models, you only need to train them on a specific subset of data. For example, if we're trying to train on stocks and business data for a specific group, we can train a smaller business model on that and we have a lots, uh, faster query talk. Also, running a small language model locally on the device means that you don't have, you barely have any latency between data coming in and a response coming out.
Again, this is really important when a 10th of a second can fluctuate a stock price massively and it's really important to be on top of the game and have really fast and real time responses. And thirdly, as on edge natural language processing, nearly any natural language task or transformation process on the edge can be handled with a small language model. Local processing means that no data is really being transmitted to the cloud and it can all be handled on the edge, which is really good because not only is it a lot faster, but it also keeps your data really secure.
For example, if you're dealing with a confidential, uh, if you're dealing with confidential data on the edge or you're dealing with something like a home assistance where you're tracking users' daily habits, you might not want that data to be all be sent back to a data center or sent in the cloud where it can all be recorded. Processing all the data on the edge means that no data needs to ever be sent over the internet and it can all be really secure. It's also really easy to connect to and it's really easy to control and easier implementation through small language models because you don't need a central hub to connect to to receive a command every time.
And fourth is in specific industries like medicine for example. Again, remember that domain specificity means that small language models are very effective in a one specific area in the medical industry. For example, a small language model could be, uh, created for diagnosing a cold or specifically trained in something for noticing any problems with an e surgery.
This means that for things like real-time assessment in the emergency room, small language models are great, but secondly in things like fast reactions in surgeries where again we can't afford any delay, small language models running on the edge and locally is really great. And finally, we can apply small language models not just in a hospital but also in telemedicine because small language models are able to be personalized to a large degree to the person that they're made for. This means that a small language model is able to personalize the diagnoses for each person.
But finally, and perhaps the most important point is that again, the small language models do not need internet connection to be ran. Meaning that you don't need to trans, you don't need to transfer any sensitive or confidential healthcare data of your patient. Next is machinery responses.
Domain specificness means that you can tailor a small English model to a specific machine reading or sensor and get really good real time fast decision making That's also offline. For example, if you wanted to put a small language model or something like that in a solar or wind farm, it's offline. It can make a decision in a split second without needing to contact a central hub.
And lastly is personal assistance because small language models have been designed with the goal of keeping them as small as possible in mind, this means that you can often host a small language model without much processing power at all. You could host it on a phone or you could even host it on a single chip or a micro, but this means that you're able to develop a lot more personal assistance without needing to connect to the cloud or without needing large amounts of processing power or putting one in a computer for example. So we talked a lot about the applications of small language models, but let's look at some of the actual small language models that exists and how they're working.
As you can see on this chart up here, there are a lot of small language models out there and this chart only scratches the surface but discusses some of the more mainstream ones. The first three on the top, this still the Berta and Tiny Bert are all SM are all, are all small language models. Based on the Bert model, which was an early primitive language model developed in 2018 by Google, these smaller language models have improved training and optimized databases.
These models are really, are significantly smaller than the main Bert model, but they're able to produce similarly accurate results or even in some cases they're, they're able to produce more accurate results than the old model that was a lot bigger that these models were based upon. Flan is another model that was developed at Google in 2021, which is a tuning model that excels in few shot learning, which is when a conversation is only a couple of prompts long and you're training it off conversations that are maybe three or four prompts long, that's where flan comes in and offers really good training sets. GPT and Neo and GBTJ are training in inference models that are based upon GPT-3 model and they're aiming to emulate GT three.
And mixed role is a unique model that operates on a mixture of experts and a mixture of experts or MOE model. There are several specialist submodels inside of this small language model itself. You can kind of think of these as a smaller small language model or a sub SLM that's operating to help this bigger SLM, which is mixed role work together.
These smaller s SLMs are each even more specified than drill. These smaller small language models are specialist submodules that are known as experts, which are each designed to handle a specific task or handle a specific piece of data within the larger model. The key component of mixed drill and the key component of any mixture of expert model is the gating network or router that the decides that once a piece of data comes in, the router is able to look at all of the options available to it.
It can see that it has a bunch of smaller sub expert uh, models that are available to use. And once the data has come in, the router can decide, oh, uh, let's see, maybe expert number one is best equipped to handle this data and the next data goes to expert number four and those experts specifically process the data and send it output. The key part of a mixtures of experts model is the gating network that decides which expert should be best equipped to handle a given input based on the characteristics of the input.
Next is FI two. FI two is a groundbreaking small language model develop that Microsoft, unlike most S SLMs FI two is actually designed for broad applications and a wide range of queries as opposed to being domain specific. However, FI two is still really small.
7 billion parameters and it can still hold its own inaccuracy and speed tests when compared to much larger language models that are two times or up to 50 times its same size. We'll see that in the following slides. ORCA two is also a model developed on Microsoft and it's based on LAMA two and has around 10 billion parameters.
ORCA two is specifically designed to test advanced reasoning capabilities in settings that I don't have any problems before or known as zero shot settings. A zero shot setting for example, would be if I open chat GBT in a new window and created a new conversation with it and ask it a question. There's no previous information that the road that the bot has, it's just going off of its pre-trained dataset.
ORCA two has been reported to have uh, compatible reasoning skills with models of the 10 times larger in of tests. And next is UL two R. That's a proposed training method for small language models that has the potential to cut computational costs by one half.
This will be really important as we'll see in later slides. And lastly, but definitely not the least is Gemma. Gemma is a group of open models that's from Google and it's based off Google's latest advancement in generative AI.
For Gemini, Google's latest AI advancement Gemini has been put into its Bard chatbot and Gemma is a scale down version of Gemini. In some tests, Gemma has performed similar to the Laude uh, model with roughly half the size. So let's drill down into some specific uh, aspects of these small language models specifically.
The first one is speed. The table on screen right now compares the performance of three key models, Elmo work base and distillers. We'll focus on the distiller tests relative to the Burt model.
For this example, overall distiller has been found through a bunch of testing to have the same or better performance as as the regular Burt model distiller retains 97% of the accuracy as regular Burt, even though it's 60% of the size, meaning that you're significantly cutting down on not only operating costs but also training costs. And training time. Distiller is compared on sentiment analysis, accuracy and also answer its ability to answer questions and that a bunch of tests and their results all suggests that distiller, which is a smaller language model and a distilled version of the main burn model, has performed comparably to Bert on these tasks.
The only trade off being a slight decrease in perform or the only uh, trade off being a slight trade off for efficiency, which is a lot better considering that these models are 40% smaller. This table compares the size and inference time of the Burt base and distiller models and realize that the distiller model has 60, has 66 million parameters while Burt has 110 million. Even though Burt might slightly edge it out in some specific tasks, this silver uh, still was silver, still marked a significant advancement in small language models and the optimizations that allowed 'em to overtake larger.
Next, let's discuss small language model performance. While you might be thinking that small language models perform a lot less accurately because they're smaller, even if you're able to, even if you're able to grab answers really quickly, small language models are actually able to outperform larger and broader models in their specifically trained sublimates. Now of course it wouldn't really be fair if we asked a small language model specifically trained on the medical industry how to like, uh, how to go fishing or how to create a rocket.
But if we ask this medical chat bot how to best create a how to best perform a surgery, this medical small language model can even outperform broader and larger LLMs. For example, my seven B, which is a mixture of experts model like we discussed earlier, is compared in this chart with different LAMA models which range from 7 billion parameters all the way up to 34 billion parameters In these tests, SREL has shown to outperform the LA Lama two 13 billion parameter model and offer at least comparative performance with a LAMA one 34 billion parameter model. Thus through creating smaller language models that are really specific and trained only on their domain, we're able to keep training size sets small and we're also able to keep accuracy high while simultaneously reducing the size and computational workload for these language models.
Let's now talk about FI two, which is a language model that came out of Microsoft that we discussed earlier here FI two is compared to a larger LAMA models and the Mytral model that we just talked about, FI two has less than 3 billion parameters. It's half the size of Mytral and and uh, LAMA 7 billion and it's 25 times smaller than the LAMA two 70 billion parameter model. However, as we can see in this chart, FI two offers comparative performance when tested against all of these other models and even has the highest software development score out of all of these other models.
This gives FI two in small English models as a whole. A couple of advantages as we can see in this picture FI two and its domain specificity specifically in coding for the purposes of this test shows that bigger is not always better. Rather, we should instead be concentrating our training set on what we're actually looking for in a model.
Higher accuracy runs in a similar vein because we're concentrating what we actually care about and training into the training set and training a small language model on that set data we can achieve a higher accuracy that actually outperforms models that aren't really specific to a specific task and are just broad and general. Thirdly, we can receive that small language models are a lot cheaper in resources as there's less training cost and less operating computations because just inherently the size of these models is a lot smaller. And lastly, because these size of the models are a lot smaller, we're able to be run, these models are able to be ran and we can develop these models really fast without a lot of need for expensive GP computations or electricity farms or anything like that.
So next, let's talk about some of the techniques we can use to fine tune our small language models. We alluded to this earlier, but there's still a couple of methods that we haven't talked about for tuning a small language models for now. Let's think of training a small language model as teaching an online course.
There are a couple ways we can do this. First is zero shot prompting, which would essentially be dropping a task on somebody that hasn't had any training on this task at all. Maybe we could give somebody maybe, maybe a quick overview per se, but they haven't done the actual task before and we're seeing how they can handle it without any time to prepare or any prior knowledge.
Zero shot prompting is really nice if we wanna get a feel for our small language models, basic skills or if the job is like really easy or the query is really easy that we that they can just wing it and get it done with and we don't really need any prior training. This, this method is really easy to implement because it doesn't really need any other training data. We can just load up the model and send it a query and it just responds to the prompts based on its pre-training data alone.
This method is best used for quickly testing when a model is capable of or for simple tasks and inferences that don't require a lot of specialized knowledge. And second option is one shot prompting. This would be like showing a student a single example of how to do something and that single example is really good and it's pretty correct.
Then we give them a model to imitate that this work, this strategy would work really well if we had one really good demonstration that encapsulates exactly what we need the model to learn about. And the model can replicate this really well. This means that we can teach the model pretty complex behaviors with minimal amounts of training data just by giving them one really good example.
This strategy is pretty useful if we wanna test our model's application capabilities or if we wanna do a little bit more complex analysis and we can't just give it to the language model without any prior training. And thirdly is few shot prompting. This would be like giving someone a handful of examples so they can get the general gist of the task and start to expand on it like their on their own.
It's like learning from a small but varied case study, which is enough to get confident from, but you don't really need a full course. This method allows you to give your language models a lot more complex and robust tasks, but their language models are still pretty efficient in terms of the number of examples required. This method is really good if you wanna enhance the model skills in a specific area when only a limited amount of training data is available.
And next is fine tuning In this strategy, you're essentially putting the language model through a full training program. You're giving a detailed examples and full exercises and giving a practice quizzes and it's an entire deep dive into learning. This would be best when you have the time and resources to really polish your student skills for a really complex task.
Fine tuning allows you to maximize your model performance and allows it to learn more complex and nuanced behaviors by training on a larger data set. And it's really, it's the right choice if you wanna specialize the model for dedicated tasks and you have a large amount of training data available to support the process. However, note that by using this method you're probably gonna make your model a little bit slower and increase the size of it quite significantly.
And lastly is parameter efficient. Fine tuning this method is just kind of a meta analysis of tuning. It's about really being smart with your training and focusing on the most critical aspects to improve.
It's like targeted coaching and uh, it's like targeted coaching. Instead of giving your language model the full course, instead you focus on some key points that you either wanna address or that you wanna strengthen that adapts to the language model specific needs and it can make sure that the language model is getting a lot better without overwhelming them or taking up way too many resources. This is kind of a middle ground between all of the training methods that are zero shot and all the training methods that require a lot of fine tuning and optimization.
So we talked a lot about small language models, but how are small language models gonna continue to affect our future and what can we do as look forward? First off is commoditization. As we've talked about numerous times in today's presentation.
The Kiva small language model is is they, they can be deployed nearly everywhere. They don't need to be connected to the internet and they're really small and cheap. So you can pro, you can run them on a small processor and doesn't even need an internet connection.
Small language models will become a lot more deeply embedded in everyday technologies. It'll help improve experiences across all devices, all platforms, and it'll be everywhere from smartphones even to like a household appliance like the refrigerator. It'll revolutionize our daily interaction with technology, meaning that technology can become a lot more personalized in context that wear to the user specifically making technology more intuitive and responsive.
This means the technology is not only gonna become more responsive but also more intelligent. Imagine my microwave instead of needing my inputs every time I wanna go ba warm up a pizza now it can infer my needs from previous interactions. It'll be able to realize that oh, around this time every week you seem to come here and want some popcorn.
It can do some touchless visual interactions without need for a panel and it can infer what you actually want to do. Because of this personalization and contextual awareness, it'll significantly improve the practicality and effectiveness of daily technology use all the way from our phones to our refrigerators and microwaves. Problematically.
There are a couple of challenges to generative AI at large main, namely security because models are trained on massive amounts of data, which is often sensitive or personally identified. Remember the Samsung leak where Samsung software developers inputted internal code to chat GBT and leaked that code. And second is costs.
Lot of lots of API calls every minute rack up a large cost and a company developing a training their own model has a large barrier to entry. Next is transparency as a lot of people still don't know how the model works, even including model developers. And lastly is competition.
Most internal company models don't really need to compete to be the best, but if you're working on the cutting edge of ai, these models are constantly optimizing each other out and trying to get a little bit ahead. So let's talk about a few challenges in particular. Namely first is scaling.
As models continue to grow in size and become more and more complex, the computational resources required to train and run these models increase exponentially. GPU pooling is a a proposed method to help deal with these, but these methods are still really expensive. You're still gonna need lots of GPUs and lots of electricity if you ever wanna scale up your uh, language model that's really bad as you might be able to infer for the environment.
Indeed OpenAI is trying to create its GT five model now, but they had to split it up to multiple data centers because they themselves said that if we train all the models of GPT five on the same GPUs in one data center, it would take down the electricity power grid. And secondly, security and privacy. As AI models become more and more advanced, there's a lot of risk for misuse for purposes like DeepFakes or misinformation where you can already see really fake, but also convincing videos online about political leaders speaking or about missile strikes in war zones.
And secondly is data privacy. Protecting the data of individuals whose data is used in interacting and trending with AI models is really important. These are exemplify through effects of regulations we've seen.
For example, the EU has recently passed its AI Regulation Act, which enacted sweeping bans on biometrics and face recognition technologies in an effort to protect users' privacy and protect their security going forward and ensure that these companies are collecting massive amounts of data to train AI models but also not notifying them and being aneth. So lastly, what should we really be doing? Well, a couple of things.
First is fostering collaborative development. Realize that creating AI is not just the task of a technologist or a developer. AI in its past couple of years has already revolutionized nearly every aspect of society that we're in today.
If we wanna continue developing language models and artificial intelligence, it's not just to the developers but also policy makers and educators. And secondly is education. In the similar vein, we need to make sure that the public understands the dangers of AI and how to best use it.
And third is responsible innovation emphasis must be placed under those importance of in uh, of sustainable language model development and ai. Remember that if we continue to make these models bigger and bigger, eventually we're gonna need so much power and we're gonna deal with massive blows to the environment. And lastly is effective policy.
AI is so new but it comes with, it's a variety of risks. We should follow things like the con like congressional bills that have been passed recently in the US or that use GDPR ACT or the use AI act that aim to establish a fine line between regulation and innovation. So we discussed a lot of things in today's presentation, but let's do a quick recap for us.
We started off with a high level overview of generative AI and small language models talking about current trends and generative AI and differentiating between generative AI and language models. And in language models we differentiated between large and small models. We then talked about examples and use cases of small language models.
We talked about how they can be used in things like healthcare or monitoring machine or in sales assistance. We then drilled down into a key into a few key characteristics, namely speed, performance and applicability of small language models. And we looked at some real world examples such as Distiller and Bert and we looked at fi two's capabilities when compared to a lot larger models.
We then talked about how we can be increasing use of small language models into our everyday lives in the future, but how that comes with its bunch of associated risks which we have to work together as a society to mitigate. Then on the call to action we discussed some important next steps for the field and important things to keep in mind as AI continues to develop and runs away from us. I hope you guys all learn something from today's presentation, be it from uh, small language models where they can be used, language models at large or even just artificial intelligence.
If you have any questions or wanna learn more, my contact information is displayed on the screen now. Thank you for coming to my presentation and I hope you have a rest of a good rest of the conference. Thank you.
