Putting AI into Production with Fireworks AI's Lin Qiao
Read full transcript 23 segments
-
hey friends we have some news from our hey friends we have some news from our sponsor text control they just released sponsor text control they just released sponsor text control they just released version 32 of their document processing version 32 of their document processing version 32 of their document processing Library which includes new core Library which includes new core Library which includes new core functionality such as document footnotes functionality such as document footnotes functionality such as document footnotes SVG export and more you can integrate SVG export and more you can integrate SVG export and more you can integrate document editing signing collaboration document editing signing collaboration document editing signing collaboration and PDF processing into your asp.net and PDF processing into your asp.net and PDF processing into your asp.net core and angular applications with TX core and angular applications with TX core and angular applications with TX text control these powerful libraries text control these powerful libraries text control these powerful libraries will let your developer teams focus on will let your developer teams focus on will let your developer teams focus on their core competencies while text their core competencies while text their core competencies while text control handles your digital document control handles your digital document control handles your digital document processing check out all the new processing check out all the new processing check out all the new features and see the Technologies live features and see the Technologies live features and see the Technologies live in action by visiting the demo at demos. in action by visiting the demo at demos. in action by visiting the demo at demos. text control.com that's demos. text text control.com that's demos. text text control.com that's demos. text [Music] [Music] [Music] control.com hi I'm Scott Hanselman this control.com hi I'm Scott Hanselman this control.com hi I'm Scott Hanselman this is another episode of Hansel minutes is another episode of Hansel minutes is another episode of Hansel minutes today I'm talking with Dr Lyn Chow CEO today I'm talking with Dr Lyn Chow CEO today I'm talking with Dr Lyn Chow CEO and co-founder of fireworks AI how are and co-founder of fireworks AI how are and co-founder of fireworks AI how are you today hi Scott I'm doing well thanks you today hi Scott I'm doing well thanks you today hi Scott I'm doing well thanks for having me yeah I'm very excited to for having me yeah I'm very excited to for having me yeah I'm very excited to chat with you today because I'm really chat with you today because I'm really chat with you today because I'm really trying to understand these next trying to understand these next trying to understand these next Generations of startups and medium-sized Generations of startups and medium-sized Generations of startups and medium-sized companies doing AI because I think it's companies doing AI because I think it's companies doing AI because I think it's worth noting that before you were the worth noting that before you were the worth noting that before you were the CEO and co-founder of fireworks you CEO and co-founder of fireworks you CEO and co-founder of fireworks you worked at big companies that are doing worked at big companies that are doing worked at big companies that are doing big stuff you worked at meta uh you ran big stuff you worked at meta uh you ran big stuff you worked at meta uh you ran a large team of Engineers doing AI you a large team of Engineers doing AI you a large team of Engineers doing AI you worked at LinkedIn uh earlier years uh
-
worked at LinkedIn uh earlier years uh worked at LinkedIn uh earlier years uh they must have been doing work in Ai and they must have been doing work in Ai and they must have been doing work in Ai and you also worked in research at IBM but you also worked in research at IBM but you also worked in research at IBM but rather than building the thing you built rather than building the thing you built rather than building the thing you built at any of those companies you chose to at any of those companies you chose to at any of those companies you chose to leave a giant company that's making big leave a giant company that's making big leave a giant company that's making big moves in AI referring to meta and do moves in AI referring to meta and do moves in AI referring to meta and do your own your own your own thing what what what's what's happening thing what what what's what's happening thing what what what's what's happening in one's brain how do you how do you in one's brain how do you how do you in one's brain how do you how do you decide I have an idea and I can't do it decide I have an idea and I can't do it decide I have an idea and I can't do it at Giant company I'm going to do it at at Giant company I'm going to do it at at Giant company I'm going to do it at startup h i I think I like need to be startup h i I think I like need to be startup h i I think I like need to be crazy enough to believe the direction crazy enough to believe the direction crazy enough to believe the direction I'm I'm heading on is going to work um I'm I'm heading on is going to work um I'm I'm heading on is going to work um so I think I I'm fortunate enough to be so I think I I'm fortunate enough to be so I think I I'm fortunate enough to be uh uh uh driving uh pytorch at Mana for a long driving uh pytorch at Mana for a long driving uh pytorch at Mana for a long time for uh five years and here the time for uh five years and here the time for uh five years and here the focus is kind of really driving pytorch focus is kind of really driving pytorch focus is kind of really driving pytorch from a research oriented project uh for from a research oriented project uh for from a research oriented project uh for researchers into fully full-blown in researchers into fully full-blown in researchers into fully full-blown in production and Matter's production scale production and Matter's production scale production and Matter's production scale is huge huge right so and there's a is huge huge right so and there's a is huge huge right so and there's a variety of uh applications and products variety of uh applications and products variety of uh applications and products building on top of AI on top of pytorch building on top of AI on top of pytorch building on top of AI on top of pytorch on top of infrastructure from ground up on top of infrastructure from ground up on top of infrastructure from ground up for pytorch um so so that path took us for pytorch um so so that path took us for pytorch um so so that path took us five years um and then Pi torch at the five years um and then Pi torch at the five years um and then Pi torch at the same time is a great open source project same time is a great open source project same time is a great open source project I think the team did a fantastic job uh I think the team did a fantastic job uh I think the team did a fantastic job uh driving the open source engagement driving the open source engagement driving the open source engagement through that uh I'm fortunate enough to through that uh I'm fortunate enough to through that uh I'm fortunate enough to work to see uh many other companies more work to see uh many other companies more work to see uh many other companies more tradition Enterprise companies in the tradition Enterprise companies in the tradition Enterprise companies in the industry start to uh start to get on to
-
industry start to uh start to get on to industry start to uh start to get on to this AI first this AI first this AI first move I still remember like we we work move I still remember like we we work move I still remember like we we work with Walmart we work with Disney we work with Walmart we work with Disney we work with Walmart we work with Disney we work with JN Deere um they they already using with JN Deere um they they already using with JN Deere um they they already using AI to to innovate it's very interesting AI to to innovate it's very interesting AI to to innovate it's very interesting applications like uh John Deere was uh applications like uh John Deere was uh applications like uh John Deere was uh using AI to uh to take pictures on the using AI to uh to take pictures on the using AI to uh to take pictures on the on the field there's a machine rolling on the field there's a machine rolling on the field there's a machine rolling through the field and and and they through the field and and and they through the field and and and they detect uh know random stuff on the field detect uh know random stuff on the field detect uh know random stuff on the field they like weed and they just unplug them they like weed and they just unplug them they like weed and they just unplug them to clean up the field before they plant to clean up the field before they plant to clean up the field before they plant uh things so that's is really uh things so that's is really uh things so that's is really interesting but at the same time we also interesting but at the same time we also interesting but at the same time we also heard uh a lot of feedback like they heard uh a lot of feedback like they heard uh a lot of feedback like they don't have the right software to enable don't have the right software to enable don't have the right software to enable AI they don't have have the right AI they don't have have the right AI they don't have have the right Hardware to enable AI they don't have Hardware to enable AI they don't have Hardware to enable AI they don't have even have the right team they are look even have the right team they are look even have the right team they are look more looking for us to as kind of more looking for us to as kind of more looking for us to as kind of helping them to move faster and and and helping them to move faster and and and helping them to move faster and and and and it's very clear to me the whole and it's very clear to me the whole and it's very clear to me the whole entire industry majority of them are in entire industry majority of them are in entire industry majority of them are in that state while we are in this AI first that state while we are in this AI first that state while we are in this AI first move so the big contrast between the the move so the big contrast between the the move so the big contrast between the the need here to kind of go faster use AI to need here to kind of go faster use AI to need here to kind of go faster use AI to enable deliver a lot of business impact enable deliver a lot of business impact enable deliver a lot of business impact and value to um the the lack of Supply and value to um the the lack of Supply and value to um the the lack of Supply lack of supply of a knowledge of how to lack of supply of a knowledge of how to lack of supply of a knowledge of how to do it do it do it properly and we feel like um my whole properly and we feel like um my whole properly and we feel like um my whole entire funding team we are impact driven entire funding team we are impact driven entire funding team we are impact driven we have been impact driven we have been impact driven we have been impact driven um for for a long time and we're um for for a long time and we're um for for a long time and we're gravitated towards bigger and bigger gravitated towards bigger and bigger gravitated towards bigger and bigger impact of course we can stay in Ma and
-
impact of course we can stay in Ma and impact of course we can stay in Ma and there's no question this huge impact to there's no question this huge impact to there's no question this huge impact to land but for our observation driving the land but for our observation driving the land but for our observation driving the whole entire industry going through this whole entire industry going through this whole entire industry going through this journey uh is is much much bigger so journey uh is is much much bigger so journey uh is is much much bigger so that's kind of why um we uh we start that's kind of why um we uh we start that's kind of why um we uh we start this company and uh build the business this company and uh build the business this company and uh build the business from up when when big companies and I from up when when big companies and I from up when when big companies and I think of big I mean talking like you think of big I mean talking like you think of big I mean talking like you know very large no pun intended there's know very large no pun intended there's know very large no pun intended there's very large language models and there's very large language models and there's very large language models and there's very large companies that make them uh very large companies that make them uh very large companies that make them uh you know they get a lot of press meta you know they get a lot of press meta you know they get a lot of press meta gets a lot of press when they release a gets a lot of press when they release a gets a lot of press when they release a model uh Microsoft get a lot of press model uh Microsoft get a lot of press model uh Microsoft get a lot of press they release like 53 um where does they release like 53 um where does they release like 53 um where does fireworks fit into the stack of small fireworks fit into the stack of small fireworks fit into the stack of small large all these different kinds of large all these different kinds of large all these different kinds of models are you the top of the pyramid models are you the top of the pyramid models are you the top of the pyramid that just makes them all friendly or that just makes them all friendly or that just makes them all friendly or where are you in that that that stack where are you in that that that stack where are you in that that that stack right so there are multiple ways to look right so there are multiple ways to look right so there are multiple ways to look at this um there's also when comes to J at this um there's also when comes to J at this um there's also when comes to J there are companies building there are companies building there are companies building infrastructure for training right so or infrastructure for training right so or infrastructure for training right so or inference uh so we actually can choose inference uh so we actually can choose inference uh so we actually can choose the both because pytorch has been we the both because pytorch has been we the both because pytorch has been we been building infrastructure for uh been building infrastructure for uh been building infrastructure for uh training using pytorch or inflence using training using pytorch or inflence using training using pytorch or inflence using pytorch as matter scale uh but we decide pytorch as matter scale uh but we decide pytorch as matter scale uh but we decide to focus on inference the reason is J as to focus on inference the reason is J as to focus on inference the reason is J as a technology and the fundamental a technology and the fundamental a technology and the fundamental disruption here is the content generated disruption here is the content generated disruption here is the content generated through those AI models are very similar through those AI models are very similar through those AI models are very similar or even Sur what human can generate um or even Sur what human can generate um or even Sur what human can generate um and that will create a whole slw of and that will create a whole slw of and that will create a whole slw of innovation in application and product innovation in application and product innovation in application and product space um so by Nature those Innovation
-
space um so by Nature those Innovation space um so by Nature those Innovation will be more consumer pumer or developer will be more consumer pumer or developer will be more consumer pumer or developer basing by Nature um and the matter is basing by Nature um and the matter is basing by Nature um and the matter is the biggest I would say b2c company and the biggest I would say b2c company and the biggest I would say b2c company and and and we we have seen that uh in in and and we we have seen that uh in in and and we we have seen that uh in in those kind of b2c company or like to to those kind of b2c company or like to to those kind of b2c company or like to to see or to pumer to developer space um see or to pumer to developer space um see or to pumer to developer space um your inference challenge or inference your inference challenge or inference your inference challenge or inference band is going to be proportional to your band is going to be proportional to your band is going to be proportional to your end user which the up li uh the limit is end user which the up li uh the limit is end user which the up li uh the limit is the population uh on Earth um versus the population uh on Earth um versus the population uh on Earth um versus training uh you are going to be training uh you are going to be training uh you are going to be proportional with number of researchers proportional with number of researchers proportional with number of researchers in rative speaking that is much much in rative speaking that is much much in rative speaking that is much much smaller scale so we decide to kind of smaller scale so we decide to kind of smaller scale so we decide to kind of solve a bigger scale problem and L focus solve a bigger scale problem and L focus solve a bigger scale problem and L focus on influence uh influence um the the on influence uh influence um the the on influence uh influence um the the interesting transition that home interesting transition that home interesting transition that home industry is going through before and industry is going through before and industry is going through before and after ji um there are few things again after ji um there are few things again after ji um there are few things again our target audience are app developers our target audience are app developers our target audience are app developers and product Engineers right um what and product Engineers right um what and product Engineers right um what doesn't change is all this to see to doesn't change is all this to see to doesn't change is all this to see to personer to developer application it personer to developer application it personer to developer application it need to be hyper interactive and as a need to be hyper interactive and as a need to be hyper interactive and as a matter of fact the response time start matter of fact the response time start matter of fact the response time start to kind of shrink smaller and smaller to kind of shrink smaller and smaller to kind of shrink smaller and smaller because hey that's how we get used to because hey that's how we get used to because hey that's how we get used to react faster and we like those kind of react faster and we like those kind of react faster and we like those kind of experience experience experience um and and today J is on the highest um and and today J is on the highest um and and today J is on the highest spectrum of the size and complexity and spectrum of the size and complexity and spectrum of the size and complexity and getting this fast extremely fast
-
getting this fast extremely fast getting this fast extremely fast experience is really hard and that experience is really hard and that experience is really hard and that breaks the product breaks the product breaks the product experience and the second is when those experience and the second is when those experience and the second is when those companies they hit a product Market fit companies they hit a product Market fit companies they hit a product Market fit actually hit a product Market fit is not actually hit a product Market fit is not actually hit a product Market fit is not easy once they hit a product Market fit easy once they hit a product Market fit easy once they hit a product Market fit they have to scale to large Corpus of they have to scale to large Corpus of they have to scale to large Corpus of users uh and then if the car structure users uh and then if the car structure users uh and then if the car structure is wrong they quickly bankrupt it's is wrong they quickly bankrupt it's is wrong they quickly bankrupt it's literally they have to kind of they literally they have to kind of they literally they have to kind of they didn't expect oh my God I'm running out didn't expect oh my God I'm running out didn't expect oh my God I'm running out money uh no matter who you are you have money uh no matter who you are you have money uh no matter who you are you have de pocket or not you can run run all the de pocket or not you can run run all the de pocket or not you can run run all the money pretty quickly that's common money pretty quickly that's common money pretty quickly that's common complaint we have heard about and they complaint we have heard about and they complaint we have heard about and they have to press on a break and even they have to press on a break and even they have to press on a break and even they have a viable product they cannot launch have a viable product they cannot launch have a viable product they cannot launch to have a viable business so we first to have a viable business so we first to have a viable business so we first heard these two problems and we're like heard these two problems and we're like heard these two problems and we're like hey we have we're going to solve this hey we have we're going to solve this hey we have we're going to solve this problem first be the fastest and most problem first be the fastest and most problem first be the fastest and most cost efficient inference cost efficient inference cost efficient inference engine I want I want to pause here and engine I want I want to pause here and engine I want I want to pause here and kind of see yeah I want to break down kind of see yeah I want to break down kind of see yeah I want to break down some pieces because I think that when some pieces because I think that when some pieces because I think that when the audience listens to yourself and the audience listens to yourself and the audience listens to yourself and people talk about AI they're still often people talk about AI they're still often people talk about AI they're still often trying to catch up because very there's trying to catch up because very there's trying to catch up because very there's a lot of startups that are basically a lot of startups that are basically a lot of startups that are basically just web apps that talk directly to the just web apps that talk directly to the just web apps that talk directly to the open a endpoint with and it's very open a endpoint with and it's very open a endpoint with and it's very simplistic anybody puts together a simplistic anybody puts together a simplistic anybody puts together a chatbot or a completion and you're very chatbot or a completion and you're very chatbot or a completion and you're very conscious iously at fireworks using the conscious iously at fireworks using the conscious iously at fireworks using the term inference engine making inferences term inference engine making inferences term inference engine making inferences drawing conclusions based on evidence drawing conclusions based on evidence drawing conclusions based on evidence and reasoning it's not just completions and reasoning it's not just completions and reasoning it's not just completions it's not just chats do what what why did
-
it's not just chats do what what why did it's not just chats do what what why did you focus on inference as being the most you focus on inference as being the most you focus on inference as being the most interesting problem to solve before we interesting problem to solve before we interesting problem to solve before we talk about scale and pricing right so uh talk about scale and pricing right so uh talk about scale and pricing right so uh this is the kind of choice between this is the kind of choice between this is the kind of choice between training of foundation model or training training of foundation model or training training of foundation model or training a gen model versus uh you have a gen a gen model versus uh you have a gen a gen model versus uh you have a gen model and you just ask the questions and model and you just ask the questions and model and you just ask the questions and and and then bake the the the answer and and then bake the the the answer and and then bake the the the answer into your application to power new into your application to power new into your application to power new product experiences right so so those product experiences right so so those product experiences right so so those are the two choices among these two are the two choices among these two are the two choices among these two choices we choose to uh to drive uh the choices we choose to uh to drive uh the choices we choose to uh to drive uh the platform to deliver much faster and more platform to deliver much faster and more platform to deliver much faster and more coste efficient serving as in you ask a coste efficient serving as in you ask a coste efficient serving as in you ask a question get back results that's how you question get back results that's how you question get back results that's how you inter with the chat gbt for example we inter with the chat gbt for example we inter with the chat gbt for example we basically provide the same product line basically provide the same product line basically provide the same product line to similar to uh chat GPT so we are not to similar to uh chat GPT so we are not to similar to uh chat GPT so we are not solving the training problem solving the training problem solving the training problem because it requires a different kind of because it requires a different kind of because it requires a different kind of setup where you have a lot of setup where you have a lot of setup where you have a lot of researchers you have a lot of data uh researchers you have a lot of data uh researchers you have a lot of data uh and you have a lot of compute right so and you have a lot of compute right so and you have a lot of compute right so typically what we observe is uh and that typically what we observe is uh and that typically what we observe is uh and that business is really really really business is really really really business is really really really challenging the fundamental reason is challenging the fundamental reason is challenging the fundamental reason is the Model depra depra cycle is very fast the Model depra depra cycle is very fast the Model depra depra cycle is very fast right now almost every week There's the right now almost every week There's the right now almost every week There's the new new model someone launch a new model new new model someone launch a new model new new model someone launch a new model um that is better in like certain um that is better in like certain um that is better in like certain benchmarking results and and from the benchmarking results and and from the benchmarking results and and from the model builder point of view they model builder point of view they model builder point of view they invest like month quarters or even
-
invest like month quarters or even invest like month quarters or even sometimes years of research and tons of sometimes years of research and tons of sometimes years of research and tons of computer resource and get to this great computer resource and get to this great computer resource and get to this great Innovation out but soon it got subsumed Innovation out but soon it got subsumed Innovation out but soon it got subsumed by somebody else right so the model uh by somebody else right so the model uh by somebody else right so the model uh model is deating really fast and second model is deating really fast and second model is deating really fast and second is compute DEA really fast too um is compute DEA really fast too um is compute DEA really fast too um usually in the past the life cycle of a usually in the past the life cycle of a usually in the past the life cycle of a skill usually a one hardw vendor they skill usually a one hardw vendor they skill usually a one hardw vendor they have new skill out every three years or have new skill out every three years or have new skill out every three years or two years right that's kind of the cycle two years right that's kind of the cycle two years right that's kind of the cycle and now we we're seeing kind of new and now we we're seeing kind of new and now we we're seeing kind of new skill coming skill coming skill coming out like multiple times a year multiple out like multiple times a year multiple out like multiple times a year multiple skills a year uh and then lot Several skills a year uh and then lot Several skills a year uh and then lot Several Hard vendors are like competing in the Hard vendors are like competing in the Hard vendors are like competing in the space um so like Hardware depression space um so like Hardware depression space um so like Hardware depression depression really fast also so that depression really fast also so that depression really fast also so that makes kind of the modeling space for makes kind of the modeling space for makes kind of the modeling space for training really really challenging but training really really challenging but training really really challenging but jni is the the notion the concept here jni is the the notion the concept here jni is the the notion the concept here is a following right before is a following right before is a following right before J for AI before J any AI model you have J for AI before J any AI model you have J for AI before J any AI model you have to train from ground from scratch as in to train from ground from scratch as in to train from ground from scratch as in if you want to use AI you first have if you want to use AI you first have if you want to use AI you first have need to have a machine learning research need to have a machine learning research need to have a machine learning research team to curate data curate lots of data team to curate data curate lots of data team to curate data curate lots of data and then train the model right uh and and then train the model right uh and and then train the model right uh and make sure it converges and then there make sure it converges and then there make sure it converges and then there are a lot of data issuing to kind of go are a lot of data issuing to kind of go are a lot of data issuing to kind of go back uh so that's a lot of investment back uh so that's a lot of investment back uh so that's a lot of investment now we have j j j the concept is it's a now we have j j j the concept is it's a now we have j j j the concept is it's a foundational model Foundation model as foundational model Foundation model as foundational model Foundation model as in um someone else build this Foundation
-
in um someone else build this Foundation in um someone else build this Foundation model for example MAA uh keeps releasing model for example MAA uh keeps releasing model for example MAA uh keeps releasing really great llama models you don't have really great llama models you don't have really great llama models you don't have to do the same thing you don't have to to do the same thing you don't have to to do the same thing you don't have to train from scratch you just need to use train from scratch you just need to use train from scratch you just need to use it as is or Infuse your data to align it as is or Infuse your data to align it as is or Infuse your data to align the model better towards your problem the model better towards your problem the model better towards your problem towards your data distribution and that towards your data distribution and that towards your data distribution and that alignment process is very lightweight alignment process is very lightweight alignment process is very lightweight it's called it's called it's called fine-tuning um so so that make it make fine-tuning um so so that make it make fine-tuning um so so that make it make this technology much more accessible this technology much more accessible this technology much more accessible just gen is much more accessible to a just gen is much more accessible to a just gen is much more accessible to a broader set of developers compared with broader set of developers compared with broader set of developers compared with traditional AI where a company need to traditional AI where a company need to traditional AI where a company need to be decided uh need to decide to found a be decided uh need to decide to found a be decided uh need to decide to found a substantial uh Team such reasonable size substantial uh Team such reasonable size substantial uh Team such reasonable size team uh with deep talent to be able to team uh with deep talent to be able to team uh with deep talent to be able to do this model training uh so so um so do this model training uh so so um so do this model training uh so so um so with that said uh because we focus on with that said uh because we focus on with that said uh because we focus on gen space uh and we basically leave the gen space uh and we basically leave the gen space uh and we basically leave the training problem to those uh research training problem to those uh research training problem to those uh research institution uh more more so now it's institution uh more more so now it's institution uh more more so now it's kind of uh basically two camps close kind of uh basically two camps close kind of uh basically two camps close Source versus open weights um and of Source versus open weights um and of Source versus open weights um and of course we bet on open weights heavily course we bet on open weights heavily course we bet on open weights heavily and we believe in that direction uh and we believe in that direction uh and we believe in that direction uh because I I have been kind of on on open because I I have been kind of on on open because I I have been kind of on on open source side uh through the python source side uh through the python source side uh through the python journey in the past five years and I journey in the past five years and I journey in the past five years and I deeply building uh that and then deeply building uh that and then deeply building uh that and then fireworks is just focus on enable those fireworks is just focus on enable those fireworks is just focus on enable those state of art models stateof art J models state of art models stateof art J models state of art models stateof art J models in the serving tier so enable lots of
-
in the serving tier so enable lots of in the serving tier so enable lots of chat gbt like apps and products to run chat gbt like apps and products to run chat gbt like apps and products to run on top of J through fireworks so that's on top of J through fireworks so that's on top of J through fireworks so that's our mission so if I'm a startup who's our mission so if I'm a startup who's our mission so if I'm a startup who's listening who is maybe calling Azure or listening who is maybe calling Azure or listening who is maybe calling Azure or another host or maybe they're calling uh another host or maybe they're calling uh another host or maybe they're calling uh open aim points directly and they've open aim points directly and they've open aim points directly and they've already picked their model they've gone already picked their model they've gone already picked their model they've gone in there and maybe they're using llama in there and maybe they're using llama in there and maybe they're using llama 3.2 or they're using llama you know 3 3.2 or they're using llama you know 3 3.2 or they're using llama you know 3 3.1 if they put if they switch to 3.1 if they put if they switch to 3.1 if they put if they switch to fireworks they're getting what they're fireworks they're getting what they're fireworks they're getting what they're getting lower cost massively lower cost getting lower cost massively lower cost getting lower cost massively lower cost massively higher throughput and uh a massively higher throughput and uh a massively higher throughput and uh a lower dollars per token how what is the lower dollars per token how what is the lower dollars per token how what is the secret herbs and spices that you're secret herbs and spices that you're secret herbs and spices that you're putting in between it's almost like I putting in between it's almost like I putting in between it's almost like I imagine an analogy this may not be a imagine an analogy this may not be a imagine an analogy this may not be a good analogy but like in the old days we good analogy but like in the old days we good analogy but like in the old days we would put a virtual machine on the open would put a virtual machine on the open would put a virtual machine on the open internet you'd open a port and people internet you'd open a port and people internet you'd open a port and people would call that virtual machine and would call that virtual machine and would call that virtual machine and right now you know open AI has this this right now you know open AI has this this right now you know open AI has this this API and everyone pretends to be open AI API and everyone pretends to be open AI API and everyone pretends to be open AI they are an open AP open AI compatible they are an open AP open AI compatible they are an open AP open AI compatible API but you would not today in 2024 put API but you would not today in 2024 put API but you would not today in 2024 put a virtual machine on the open internet a virtual machine on the open internet a virtual machine on the open internet you would put a caching system in front you would put a caching system in front you would put a caching system in front of it you'd put a reverse proxy you put of it you'd put a reverse proxy you put of it you'd put a reverse proxy you put a series of layers to protect it on the a series of layers to protect it on the a series of layers to protect it on the input and on the output is that what's input and on the output is that what's input and on the output is that what's happening with large language models you happening with large language models you happening with large language models you would not put a large language model on would not put a large language model on would not put a large language model on the open internet directly you would the open internet directly you would the open internet directly you would need to wrap it in safety and caching need to wrap it in safety and caching need to wrap it in safety and caching and throughput and things like that and throughput and things like that and throughput and things like that that's a very interesting analogy I that's a very interesting analogy I that's a very interesting analogy I really like it you can use really like it you can use really like it you can use that thanks for the idea uh so uh there
-
that thanks for the idea uh so uh there that thanks for the idea uh so uh there are few things right so compare with are few things right so compare with are few things right so compare with Azure or um AWS or Azure or um AWS or Azure or um AWS or gcp um so so first of all um our engine gcp um so so first of all um our engine gcp um so so first of all um our engine because we wrote pyto code right a lot because we wrote pyto code right a lot because we wrote pyto code right a lot of pyto of pyto of pyto optimization uh we're very kind of we optimization uh we're very kind of we optimization uh we're very kind of we the domain expert here and all these gen the domain expert here and all these gen the domain expert here and all these gen models are pyo models so um and we models are pyo models so um and we models are pyo models so um and we basically have a special pyto version basically have a special pyto version basically have a special pyto version for J for J for J models and here uh of course we also models and here uh of course we also models and here uh of course we also wrote a lot of low-level um like Kura wrote a lot of low-level um like Kura wrote a lot of low-level um like Kura kernels uh rocken kernels MD CA kernel kernels uh rocken kernels MD CA kernel kernels uh rocken kernels MD CA kernel MV MV MV uh but we also have a special uh but we also have a special uh but we also have a special orchestration uh and dist distributed orchestration uh and dist distributed orchestration uh and dist distributed execution engine uh that we Implement execution engine uh that we Implement execution engine uh that we Implement inside of pyal Base but which is like inside of pyal Base but which is like inside of pyal Base but which is like specialized for genni uh and we also imp specialized for genni uh and we also imp specialized for genni uh and we also imp very different level of very different level of very different level of caching uh that's why I like your caching uh that's why I like your caching uh that's why I like your analogy here uh if you can hit if have analogy here uh if you can hit if have analogy here uh if you can hit if have high cash hit ratio as in the prompts high cash hit ratio as in the prompts high cash hit ratio as in the prompts are similar to each other are similar to each other are similar to each other then you don't have to recompute every then you don't have to recompute every then you don't have to recompute every time because this is really expensive um time because this is really expensive um time because this is really expensive um so leave that aside our specialty is so leave that aside our specialty is so leave that aside our specialty is actually a different thing other than actually a different thing other than actually a different thing other than hey we have our own optimized pyo uh hey we have our own optimized pyo uh hey we have our own optimized pyo uh runtime for J um we specialize our runtime for J um we specialize our runtime for J um we specialize our observation is every person's every
-
observation is every person's every observation is every person's every developers every Enterprise their developers every Enterprise their developers every Enterprise their influence workload um is different f influence workload um is different f influence workload um is different f other everyone is different um and uh other everyone is different um and uh other everyone is different um and uh and if we treat everyone the same way and if we treat everyone the same way and if we treat everyone the same way then we leave a lot of uh latency then we leave a lot of uh latency then we leave a lot of uh latency optimization cost reduction quality optimization cost reduction quality optimization cost reduction quality improvement all on the table because we improvement all on the table because we improvement all on the table because we don't differentiate across the board one don't differentiate across the board one don't differentiate across the board one analogy I want to um use here is analogy I want to um use here is analogy I want to um use here is database has been uh a well understood database has been uh a well understood database has been uh a well understood area um it has it is declarative it has area um it has it is declarative it has area um it has it is declarative it has SQL as a language where you describe SQL as a language where you describe SQL as a language where you describe what you want to get out of the database what you want to get out of the database what you want to get out of the database but you don't describe how to do it and but you don't describe how to do it and but you don't describe how to do it and a database engine take your SQL query a database engine take your SQL query a database engine take your SQL query and then go through a process called and then go through a process called and then go through a process called query query query optimization and this qu optimization is optimization and this qu optimization is optimization and this qu optimization is going to look at what's your workload going to look at what's your workload going to look at what's your workload what's your qu look like what do you what's your qu look like what do you what's your qu look like what do you want to do and generate the most want to do and generate the most want to do and generate the most optimized plan for low uh low latency optimized plan for low uh low latency optimized plan for low uh low latency and low and low and low cost so same idea here uh everybody's cost so same idea here uh everybody's cost so same idea here uh everybody's inference workload as in kind of the the inference workload as in kind of the the inference workload as in kind of the the request you sent to chat GPT for example request you sent to chat GPT for example request you sent to chat GPT for example it's different from each other really it's different from each other really it's different from each other really depends on use case depends on your depends on use case depends on your depends on use case depends on your internal data sets and so on uh we internal data sets and so on uh we internal data sets and so on uh we can heavily personalize or optimize can heavily personalize or optimize can heavily personalize or optimize infent towards your
-
infent towards your infent towards your workload uh and that's called fireworks workload uh and that's called fireworks workload uh and that's called fireworks Optimizer fire Optimizer fire Optimizer fire Optimizer uh similar idea to query Optimizer uh similar idea to query Optimizer uh similar idea to query optimizer for databases and that makes optimizer for databases and that makes optimizer for databases and that makes us stand out for everybody else is um we us stand out for everybody else is um we us stand out for everybody else is um we do special like customization and do special like customization and do special like customization and personalization to the extreme uh and by personalization to the extreme uh and by personalization to the extreme uh and by using file Optimizer uh we let you using file Optimizer uh we let you using file Optimizer uh we let you choose so you need to imagine um here choose so you need to imagine um here choose so you need to imagine um here the design space is the design space is the design space is multi-dimensional um across quality multi-dimensional um across quality multi-dimensional um across quality model accuracy quality latency and cost model accuracy quality latency and cost model accuracy quality latency and cost it's a three-dimensional uh tradeoff it's a three-dimensional uh tradeoff it's a three-dimensional uh tradeoff space space space uh and there's a curve um you can uh you uh and there's a curve um you can uh you uh and there's a curve um you can uh you can get to and then everyone want to can get to and then everyone want to can get to and then everyone want to choose a different point in the curve to choose a different point in the curve to choose a different point in the curve to optimize for we expose this curve to optimize for we expose this curve to optimize for we expose this curve to them but everyone's curve is also them but everyone's curve is also them but everyone's curve is also different because your workloads are different because your workloads are different because your workloads are very unique and distinct from each other very unique and distinct from each other very unique and distinct from each other so we basically present uh this so we basically present uh this so we basically present uh this personalized curve to everyone and they personalized curve to everyone and they personalized curve to everyone and they can pick and choose which point uh they can pick and choose which point uh they can pick and choose which point uh they want to land and that's far want to land and that's far want to land and that's far Optimizer the uh that reminds me of the Optimizer the uh that reminds me of the Optimizer the uh that reminds me of the old software engineering joke that you old software engineering joke that you old software engineering joke that you can have it good fast or cheap pick two we want you to pick three yeah you two we want you to pick three yeah you want to pick I would love it to be high want to pick I would love it to be high want to pick I would love it to be high quality lowc cost and highly efficient quality lowc cost and highly efficient quality lowc cost and highly efficient but by exposing that curve you're but by exposing that curve you're but by exposing that curve you're allowing them that level of flexibility allowing them that level of flexibility allowing them that level of flexibility because um you know right now a lot of because um you know right now a lot of because um you know right now a lot of you know I think we as an industry are you know I think we as an industry are you know I think we as an industry are still discovering what the UI should be
-
still discovering what the UI should be still discovering what the UI should be for this for the customer if you look at for this for the customer if you look at for this for the customer if you look at like Bing chat they say precise versus like Bing chat they say precise versus like Bing chat they say precise versus creative and you know for a developer creative and you know for a developer creative and you know for a developer they might have a temperature slider and they might have a temperature slider and they might have a temperature slider and if you go to the open AI uh you know if you go to the open AI uh you know if you go to the open AI uh you know playground they allow the temperature to playground they allow the temperature to playground they allow the temperature to go to two which one could argue has zero go to two which one could argue has zero go to two which one could argue has zero value but if you go to Azure open AI value but if you go to Azure open AI value but if you go to Azure open AI they don't even allow temperatures above they don't even allow temperatures above they don't even allow temperatures above 0.8 because they've made a declaration 0.8 because they've made a declaration 0.8 because they've made a declaration that irresponsible things happen at that irresponsible things happen at that irresponsible things happen at temperatures over over one so when I temperatures over over one so when I temperatures over over one so when I went into the fireworks uh devel veler went into the fireworks uh devel veler went into the fireworks uh devel veler playground where folks can go and take a playground where folks can go and take a playground where folks can go and take a look at fireworks. and explore you know look at fireworks. and explore you know look at fireworks. and explore you know you expose a lot of those sliders and you expose a lot of those sliders and you expose a lot of those sliders and those you know knobs and allow people to those you know knobs and allow people to those you know knobs and allow people to uh to experiment like anyone but you do uh to experiment like anyone but you do uh to experiment like anyone but you do allow for potentially high temperatures allow for potentially high temperatures allow for potentially high temperatures and uh I don't know if a high and uh I don't know if a high and uh I don't know if a high temperature necessarily is a creative I temperature necessarily is a creative I temperature necessarily is a creative I don't know why we decided that that was don't know why we decided that that was don't know why we decided that that was the creativity the creativity the creativity axis yeah so like it really depends on axis yeah so like it really depends on axis yeah so like it really depends on your use case and application right so your use case and application right so your use case and application right so for some use case for example for some use case for example for some use case for example classification classification classification um and they use uh LMS to classify hey um and they use uh LMS to classify hey um and they use uh LMS to classify hey what kind of um what what kind of what kind of um what what kind of what kind of um what what kind of whether um like for example product whether um like for example product whether um like for example product catalog right what kind of category catalog right what kind of category catalog right what kind of category should that fit into uh and they need a should that fit into uh and they need a should that fit into uh and they need a very like you don't want to holis very like you don't want to holis very like you don't want to holis inate you don't want to kind of build inate you don't want to kind of build inate you don't want to kind of build based on probability you want to kind of based on probability you want to kind of based on probability you want to kind of this very uh very accurate um
-
this very uh very accurate um this very uh very accurate um deterministic Behavior so you kind of deterministic Behavior so you kind of deterministic Behavior so you kind of temperature zero no no kind of temperature zero no no kind of temperature zero no no kind of creativity is needed here um but creativity is needed here um but creativity is needed here um but sometimes you want to we have seen kind sometimes you want to we have seen kind sometimes you want to we have seen kind of use cases of use cases of use cases like uh you want to kind of paraphrase like uh you want to kind of paraphrase like uh you want to kind of paraphrase uh some kind of for email right so it uh some kind of for email right so it uh some kind of for email right so it depends on different role you want to depends on different role you want to depends on different role you want to play uh if uh if you want to make it play uh if uh if you want to make it play uh if uh if you want to make it more interesting uh more unique or more more interesting uh more unique or more more interesting uh more unique or more busy or professional uh then you kind of busy or professional uh then you kind of busy or professional uh then you kind of the temperature does play a role in the temperature does play a role in the temperature does play a role in adjusting uh the that I'll come here adjusting uh the that I'll come here adjusting uh the that I'll come here interesting so you you call out on the interesting so you you call out on the interesting so you you call out on the site that it is the best way to go from site that it is the best way to go from site that it is the best way to go from prototype to production what are things prototype to production what are things prototype to production what are things that you think people are not thinking that you think people are not thinking that you think people are not thinking about when they go into production with about when they go into production with about when they go into production with language models my first thought is language models my first thought is language models my first thought is responsibility to prevent problematic responsibility to prevent problematic responsibility to prevent problematic things coming out of the models uh when things coming out of the models uh when things coming out of the models uh when problematic things are maybe sent into problematic things are maybe sent into problematic things are maybe sent into them um what kind of safety and where's them um what kind of safety and where's them um what kind of safety and where's your stance on when it comes to what we your stance on when it comes to what we your stance on when it comes to what we would call responsible or you know low would call responsible or you know low would call responsible or you know low bias or bias free free uh bias or bias free free uh bias or bias free free uh AIS that's a really big question I think um that comes to think um that comes to actually like different model providers actually like different model providers actually like different model providers so for example we work with many model so for example we work with many model so for example we work with many model providers some are even proprietary providers some are even proprietary providers some are even proprietary models hosted on models hosted on models hosted on fireworks um everybody is putting a lot fireworks um everybody is putting a lot fireworks um everybody is putting a lot of effort in safety features uh like of effort in safety features uh like of effort in safety features uh like recently MAA announced um launch Lama
-
recently MAA announced um launch Lama recently MAA announced um launch Lama 3.2 where their launch partner um and I 3.2 where their launch partner um and I 3.2 where their launch partner um and I I gave a presentation in matter connect I gave a presentation in matter connect I gave a presentation in matter connect um in the day of launch uh and that um in the day of launch uh and that um in the day of launch uh and that there's a backstage and we we just chat there's a backstage and we we just chat there's a backstage and we we just chat about uh different things and and and about uh different things and and and about uh different things and and and clearly the conversation becomes quickly clearly the conversation becomes quickly clearly the conversation becomes quickly becomes how to ensure safety right becomes how to ensure safety right becomes how to ensure safety right there's a lot of effort mattera put into there's a lot of effort mattera put into there's a lot of effort mattera put into the model itself uh because they don't the model itself uh because they don't the model itself uh because they don't have a hosted API they can control and have a hosted API they can control and have a hosted API they can control and kind of model need to really have the kind of model need to really have the kind of model need to really have the god rail god rail god rail um so so so it goes into the kind of um so so so it goes into the kind of um so so so it goes into the kind of post training doing lot of work to make post training doing lot of work to make post training doing lot of work to make sure the model just doesn't spit out uh sure the model just doesn't spit out uh sure the model just doesn't spit out uh bad bad bad things um yeah so I think that is kind things um yeah so I think that is kind things um yeah so I think that is kind of one thing is kind of the model of one thing is kind of the model of one thing is kind of the model provider need to do a lot of work and I provider need to do a lot of work and I provider need to do a lot of work and I do see in across industry people are do see in across industry people are do see in across industry people are holding a higher and higher bar the holding a higher and higher bar the holding a higher and higher bar the second is I will say like safety second is I will say like safety second is I will say like safety security in the Enterprise for different security in the Enterprise for different security in the Enterprise for different ENT Enterprise for different industry ENT Enterprise for different industry ENT Enterprise for different industry segment means different segment means different segment means different things um so often time um like the things um so often time um like the things um so often time um like the industry themsel has their own industry themsel has their own industry themsel has their own guideline so that just means you have to guideline so that just means you have to guideline so that just means you have to customize the model to follow your customize the model to follow your customize the model to follow your guideline better um and that's where guideline better um and that's where guideline better um and that's where kind of goes back to fire Optimizer kind of goes back to fire Optimizer kind of goes back to fire Optimizer right so we don't believe in oneid fit right so we don't believe in oneid fit right so we don't believe in oneid fit all uh and there's a lot of like we do all uh and there's a lot of like we do all uh and there's a lot of like we do see the future trend is everyone can
-
see the future trend is everyone can see the future trend is everyone can easily customize the model to fit in not easily customize the model to fit in not easily customize the model to fit in not just make it better quality for for for just make it better quality for for for just make it better quality for for for your workload but also for the for the your workload but also for the for the your workload but also for the for the for the safety reasons right so you may for the safety reasons right so you may for the safety reasons right so you may have a different safety guideline from have a different safety guideline from have a different safety guideline from those model builder because it's similar those model builder because it's similar those model builder because it's similar to model when you build a model uh you to model when you build a model uh you to model when you build a model uh you make assumptions what this model is good make assumptions what this model is good make assumptions what this model is good for right similarly when you build a for right similarly when you build a for right similarly when you build a safety uh guard re you make assumptions safety uh guard re you make assumptions safety uh guard re you make assumptions what the safety guard re should look what the safety guard re should look what the safety guard re should look like but different region for example like but different region for example like but different region for example has different has different has different uh different bars and different needs uh different bars and different needs uh different bars and different needs different industry segment also so I different industry segment also so I different industry segment also so I believe in happy customization so you believe in happy customization so you believe in happy customization so you can add uh adjust um and enhance your can add uh adjust um and enhance your can add uh adjust um and enhance your safety guideline safety guideline safety guideline accordingly hey friends in the interest accordingly hey friends in the interest accordingly hey friends in the interest of me learning more about Ai and helping of me learning more about Ai and helping of me learning more about Ai and helping you learn more about AI for the next you learn more about AI for the next you learn more about AI for the next couple of episodes we're going to have couple of episodes we're going to have couple of episodes we're going to have Lander vandoran from Qualcomm help me Lander vandoran from Qualcomm help me Lander vandoran from Qualcomm help me understand understand understand AI what is an npu and why is it AI what is an npu and why is it AI what is an npu and why is it important to today uh so the npu stands important to today uh so the npu stands important to today uh so the npu stands for uh neural processing units and it's for uh neural processing units and it's for uh neural processing units and it's basically your AI accelerator and so it basically your AI accelerator and so it basically your AI accelerator and so it can execute AI operations these large can execute AI operations these large can execute AI operations these large language models language models language models or all sorts of different generative AI or all sorts of different generative AI or all sorts of different generative AI models in a much more efficient way than models in a much more efficient way than models in a much more efficient way than a CPU or a GPU can so just to give you a CPU or a GPU can so just to give you a CPU or a GPU can so just to give you an example with Microsoft's recall right an example with Microsoft's recall right an example with Microsoft's recall right you run that continuously in the ground you run that continuously in the ground you run that continuously in the ground running it on an npu takes about 30
-
running it on an npu takes about 30 running it on an npu takes about 30 minutes of overall life battery lifetime minutes of overall life battery lifetime minutes of overall life battery lifetime if you were to run it on a GPU it would if you were to run it on a GPU it would if you were to run it on a GPU it would take six and a half hours right so it's take six and a half hours right so it's take six and a half hours right so it's all about Energy Efficiency very all about Energy Efficiency very all about Energy Efficiency very specifically dedicated to a single task specifically dedicated to a single task specifically dedicated to a single task which is AI in this case very cool so which is AI in this case very cool so which is AI in this case very cool so I'm going to go learn more about npus I'm going to go learn more about npus I'm going to go learn more about npus and how they can do things more cheaply and how they can do things more cheaply and how they can do things more cheaply from a battery perspective more from a battery perspective more from a battery perspective more efficiently and why I might want a efficiently and why I might want a efficiently and why I might want a Snapdragon processor in my next PC Snapdragon processor in my next PC Snapdragon processor in my next PC exactly thank you and folks can check exactly thank you and folks can check exactly thank you and folks can check out more at out more at out more at qualcomm.com developer uh to download qualcomm.com developer uh to download qualcomm.com developer uh to download those things and learn more those things and learn more those things and learn more today does fire Optimizer is is known today does fire Optimizer is is known today does fire Optimizer is is known for this speculative decoding does in for this speculative decoding does in for this speculative decoding does in running things in in parallel and you running things in in parallel and you running things in in parallel and you you're getting a really high hit rate you're getting a really high hit rate you're getting a really high hit rate and forgive me if this is a dumb and forgive me if this is a dumb and forgive me if this is a dumb question but does that have anything to question but does that have anything to question but does that have anything to do with or does that guarantee a higher do with or does that guarantee a higher do with or does that guarantee a higher quality inference or a less problematic quality inference or a less problematic quality inference or a less problematic one if you're doing that kind of uh uh one if you're doing that kind of uh uh one if you're doing that kind of uh uh speculative decoding right so specul speculative decoding right so specul speculative decoding right so specul decoding in a nutshell is basically you decoding in a nutshell is basically you decoding in a nutshell is basically you think about a pair of models this is think about a pair of models this is think about a pair of models this is just analogy right so sure sure a pair just analogy right so sure sure a pair just analogy right so sure sure a pair of models a small big right uh and then of models a small big right uh and then of models a small big right uh and then by running them together you get the by running them together you get the by running them together you get the latency of the small model you get the latency of the small model you get the latency of the small model you get the quality of the big model basically you quality of the big model basically you quality of the big model basically you get the Battle of both world right so get the Battle of both world right so get the Battle of both world right so but for this to be effective but for this to be effective but for this to be effective the the small model need to answer most the the small model need to answer most the the small model need to answer most of the questions right and then
-
of the questions right and then of the questions right and then sometimes like some of course it's the sometimes like some of course it's the sometimes like some of course it's the grity is not that question level I'm grity is not that question level I'm grity is not that question level I'm just drawing analogy here uh but just drawing analogy here uh but just drawing analogy here uh but sometimes if the small model is not good sometimes if the small model is not good sometimes if the small model is not good at uh giving answer correct answers at uh giving answer correct answers at uh giving answer correct answers aligning with big model then you most of aligning with big model then you most of aligning with big model then you most of time it it becomes a big model answering time it it becomes a big model answering time it it becomes a big model answering the question and then you don't get the the question and then you don't get the the question and then you don't get the small model latency right so basically small model latency right so basically small model latency right so basically it's a small model hit ratio to be very it's a small model hit ratio to be very it's a small model hit ratio to be very high and then how to increase that right high and then how to increase that right high and then how to increase that right so basically the small model and big so basically the small model and big so basically the small model and big model they need to align the more model they need to align the more model they need to align the more aligned they are um then uh the more aligned they are um then uh the more aligned they are um then uh the more likely you'll get the benefit from both likely you'll get the benefit from both likely you'll get the benefit from both but alignment is not done in in vacuum but alignment is not done in in vacuum but alignment is not done in in vacuum this alignment need to happen in tandem this alignment need to happen in tandem this alignment need to happen in tandem uh with your inference workload so both uh with your inference workload so both uh with your inference workload so both need to align with your inference need to align with your inference need to align with your inference workload basically for specific narrowly workload basically for specific narrowly workload basically for specific narrowly defined problem you want to solve um defined problem you want to solve um defined problem you want to solve um both are good at answering those and both are good at answering those and both are good at answering those and then you get the benefit of both you get then you get the benefit of both you get then you get the benefit of both you get the lower leny from small model but U the lower leny from small model but U the lower leny from small model but U higher quality from big model so okay higher quality from big model so okay higher quality from big model so okay that makes a lot of sense Optimizer is that makes a lot of sense Optimizer is that makes a lot of sense Optimizer is the provide automatic way to do that the provide automatic way to do that the provide automatic way to do that alignment uh so basically you don't to alignment uh so basically you don't to alignment uh so basically you don't to even worry about that we take care of even worry about that we take care of even worry about that we take care of that for you uh as long as kind of you that for you uh as long as kind of you that for you uh as long as kind of you use fireworks inference engine yeah I use fireworks inference engine yeah I use fireworks inference engine yeah I know you you've probably learned this know you you've probably learned this know you you've probably learned this about me in just the short time we spent about me in just the short time we spent about me in just the short time we spent together that I'm all about analogies uh together that I'm all about analogies uh together that I'm all about analogies uh because if I can understand something in because if I can understand something in because if I can understand something in another world then maybe I'll be able to another world then maybe I'll be able to another world then maybe I'll be able to understand this new world of AI this understand this new world of AI this understand this new world of AI this reminds me of pees and ecores
-
reminds me of pees and ecores reminds me of pees and ecores performance cores and efficiency cores performance cores and efficiency cores performance cores and efficiency cores and the idea of asymmetrical design on a and the idea of asymmetrical design on a and the idea of asymmetrical design on a PC where you've got these processors PC where you've got these processors PC where you've got these processors that are going to use lots of power and that are going to use lots of power and that are going to use lots of power and lots of wattage and you're going to get lots of wattage and you're going to get lots of wattage and you're going to get great fast throughput but I'm just great fast throughput but I'm just great fast throughput but I'm just checking email so I'll just do that work checking email so I'll just do that work checking email so I'll just do that work on the ecor and with asymmetrical design on the ecor and with asymmetrical design on the ecor and with asymmetrical design you get effectively if you do it right you get effectively if you do it right you get effectively if you do it right the best of both worlds and it feels the best of both worlds and it feels the best of both worlds and it feels like you are describing a kind of like you are describing a kind of like you are describing a kind of asymmetrical design in this uh in this asymmetrical design in this uh in this asymmetrical design in this uh in this speculative decoding right but this speculative decoding right but this speculative decoding right but this asymmetric design we can it's also asymmetric design we can it's also asymmetric design we can it's also adaptable it will adapt so imagine your adaptable it will adapt so imagine your adaptable it will adapt so imagine your um your e core and Edge core emission um your e core and Edge core emission um your e core and Edge core emission yeah p p cores and E cores efficiency yeah p p cores and E cores efficiency yeah p p cores and E cores efficiency versus performance so they are going to versus performance so they are going to versus performance so they are going to adapt based on what kind of workload you adapt based on what kind of workload you adapt based on what kind of workload you send to it right it can adjust itself send to it right it can adjust itself send to it right it can adjust itself right of course Hardware is impossible right of course Hardware is impossible right of course Hardware is impossible but kind of imagine if that happens but kind of imagine if that happens but kind of imagine if that happens that's a killer future so so there yeah that's a killer future so so there yeah that's a killer future so so there yeah that's what we do right and the magic on that's what we do right and the magic on that's what we do right and the magic on the in my analogy is the konel scheduler the in my analogy is the konel scheduler the in my analogy is the konel scheduler and the magic in your analogy is the and the magic in your analogy is the and the magic in your analogy is the fire Optimizer right exactly right and I fire Optimizer right exactly right and I fire Optimizer right exactly right and I like that you kept saying and then you like that you kept saying and then you like that you kept saying and then you don't have to worry about it because don't have to worry about it because don't have to worry about it because this is the thing right you're a this is the thing right you're a this is the thing right you're a customer service representative you're customer service representative you're customer service representative you're trying to help people with customer trying to help people with customer trying to help people with customer service chat problems about your service chat problems about your service chat problems about your specific domain and you've selected a specific domain and you've selected a specific domain and you've selected a very large language model that knows how very large language model that knows how very large language model that knows how to make make songs about Star Wars in to make make songs about Star Wars in to make make songs about Star Wars in the form of Shakespeare and none of that the form of Shakespeare and none of that the form of Shakespeare and none of that has anything to do with the business has anything to do with the business has anything to do with the business impact that you're trying to deliver and
-
impact that you're trying to deliver and impact that you're trying to deliver and that's why you want the quality of the that's why you want the quality of the that's why you want the quality of the large language model but you want to large language model but you want to large language model but you want to ignore all of that extra Corpus that ignore all of that extra Corpus that ignore all of that extra Corpus that unnecessary information and focus on the unnecessary information and focus on the unnecessary information and focus on the fine-tuned goal of your business that's fine-tuned goal of your business that's fine-tuned goal of your business that's right so that's why I think at the right so that's why I think at the right so that's why I think at the beginning you ask a question right beginning you ask a question right beginning you ask a question right about small model or big model right I'm about small model or big model right I'm about small model or big model right I'm a firm believer that in kind of in the a firm believer that in kind of in the a firm believer that in kind of in the business world when we solve business business world when we solve business business world when we solve business problems we absolutely are going to a problems we absolutely are going to a problems we absolutely are going to a direction of small models and as a direction of small models and as a direction of small models and as a matter of fact it's going to go smaller matter of fact it's going to go smaller matter of fact it's going to go smaller and smaller because a business task is a and smaller because a business task is a and smaller because a business task is a constraint problem it's not open-ended constraint problem it's not open-ended constraint problem it's not open-ended let's solve like let's solve a math like let's solve like let's solve a math like let's solve like let's solve a math like very hard math problem or anything uh so very hard math problem or anything uh so very hard math problem or anything uh so it's very specific um and exactly as you it's very specific um and exactly as you it's very specific um and exactly as you said when the problem is very specific said when the problem is very specific said when the problem is very specific Nar defined you can throw away other Nar defined you can throw away other Nar defined you can throw away other things um and you don't need like um things um and you don't need like um things um and you don't need like um billion parameter JM models um so uh billion parameter JM models um so uh billion parameter JM models um so uh like trillion parameter um gen models so like trillion parameter um gen models so like trillion parameter um gen models so instead like billion parameter uh or instead like billion parameter uh or instead like billion parameter uh or even sub building parameter gen models even sub building parameter gen models even sub building parameter gen models will be able to do equally good job for will be able to do equally good job for will be able to do equally good job for a specific narrow defined task uh and a specific narrow defined task uh and a specific narrow defined task uh and then the rest will become customization then the rest will become customization then the rest will become customization small model plus customization is going small model plus customization is going small model plus customization is going to be the Killer M you know I we we as a to be the Killer M you know I we we as a to be the Killer M you know I we we as a as a culture in the news cycle as we as a culture in the news cycle as we as a culture in the news cycle as we watch the personalities for lack of a watch the personalities for lack of a watch the personalities for lack of a better word at open AI have their
-
better word at open AI have their better word at open AI have their personalities we have this fascination personalities we have this fascination personalities we have this fascination with the largest models and meta has with the largest models and meta has with the largest models and meta has been leaning in on saying massive been leaning in on saying massive been leaning in on saying massive context window unlimited context you context window unlimited context you context window unlimited context you know paste your entire life into a know paste your entire life into a know paste your entire life into a that's your context window billion that's your context window billion that's your context window billion parameters but I think that the really parameters but I think that the really parameters but I think that the really interesting thing is going to be What interesting thing is going to be What interesting thing is going to be What Can I Do privately on my own machine on Can I Do privately on my own machine on Can I Do privately on my own machine on the pocket supercomputer I have in my in the pocket supercomputer I have in my in the pocket supercomputer I have in my in my phone on the npu that I'll have in my my phone on the npu that I'll have in my my phone on the npu that I'll have in my laptop I I am disappointed in our laptop I I am disappointed in our laptop I I am disappointed in our fascination with very large language fascination with very large language fascination with very large language models that we pretend our AGI when I models that we pretend our AGI when I models that we pretend our AGI when I think really interesting work happens in think really interesting work happens in think really interesting work happens in a group of small models talking to each a group of small models talking to each a group of small models talking to each other and voting on things that's right other and voting on things that's right other and voting on things that's right that's exactly that's what happening in that's exactly that's what happening in that's exactly that's what happening in the industry right so uh we have seen the industry right so uh we have seen the industry right so uh we have seen the power um of small model can do big the power um of small model can do big the power um of small model can do big things um and uh and again a small model things um and uh and again a small model things um and uh and again a small model combin with your data asset and combin with your data asset and combin with your data asset and basically enable the developers and basically enable the developers and basically enable the developers and Enterprise to have a path to convert Enterprise to have a path to convert Enterprise to have a path to convert their data asset into a their data asset into a their data asset into a assets uh and channel that value into assets uh and channel that value into assets uh and channel that value into their apps and products so that's the their apps and products so that's the their apps and products so that's the transition the whole entire industry is transition the whole entire industry is transition the whole entire industry is going through yeah well this has been going through yeah well this has been going through yeah well this has been really really educational I feel like really really educational I feel like really really educational I feel like I've got even more things to dig into I've got even more things to dig into I've got even more things to dig into I'm going to continue to play at the I'm going to continue to play at the I'm going to continue to play at the fire fireworks AI dashboard folks can fire fireworks AI dashboard folks can fire fireworks AI dashboard folks can check out fireworks. aai just click try check out fireworks. aai just click try check out fireworks. aai just click try now log in and you can explore they've
-
now log in and you can explore they've now log in and you can explore they've got all the latest llama 3.2 models got all the latest llama 3.2 models got all the latest llama 3.2 models available now you can read the blog post available now you can read the blog post available now you can read the blog post and uh hit get started right there on and uh hit get started right there on and uh hit get started right there on the home screen and you'll be playing the home screen and you'll be playing the home screen and you'll be playing with the platform very very quickly with the platform very very quickly with the platform very very quickly thank you so much uh Dr lyncha for thank you so much uh Dr lyncha for thank you so much uh Dr lyncha for chatting with me today oh definitely chatting with me today oh definitely chatting with me today oh definitely thanks for having me Scott this has been thanks for having me Scott this has been thanks for having me Scott this has been another episode of Hansel minutes and another episode of Hansel minutes and another episode of Hansel minutes and we'll see you again next week we'll see you again next week we'll see you again next week [Music]
Summary
This transcript features a discussion about Fireworks AI, a startup founded by Dr. Lyn Chow, who previously held leadership roles in AI at Meta and LinkedIn. The conversation explores the transition from large-scale corporate AI to building a startup, emphasizing the drive and belief needed to pursue a novel direction. The key takeaway is the importance of embracing bold ideas and the passion required to bring them to fruition in the competitive AI landscape.