Small Models, Big Results: Training a Finance Agent for Under $500 — Charles Dickens, Snorkel AI
Read full transcript 13 segments
-
Hello everyone, thank you for Hello everyone, thank you for joining. I'm Charles joining. I'm Charles , I'm a research , I'm a research , I'm a research scientist at Snorkel AI. scientist at Snorkel AI. Today I'll talk Today I'll talk about how we managed to about how we managed to about how we managed to make a model with 4 make a model with 4 make a model with 4 billion parameters billion parameters billion parameters outperform its outperform its outperform its version with 235 version with 235 version with 235 billion. This is a billion. This is a billion. This is a specific result specific result specific result in a simulated in a simulated in a simulated financial financial financial environment, but I environment, but I environment, but I believe there are general believe there are general believe there are general conclusions worth conclusions worth conclusions worth learning. First, learning. First, learning. First, for enterprise for enterprise for enterprise AI tasks, AI tasks, AI tasks, reliability and reliability and reliability and specialization may specialization may specialization may outweigh the sheer outweigh the sheer outweigh the sheer number of parameters number of parameters . This is possible thanks to . This is possible thanks to . This is possible thanks to high-quality data high-quality data high-quality data that truly that truly that truly reflects the domain and reflects the domain and reflects the domain and task distribution task distribution task distribution you are targeting. you are targeting. This could lead to This could lead to smaller smaller smaller models that models that models that achieve the level of achieve the level of achieve the level of advanced models at a advanced models at a advanced models at a much lower price. much lower price. Before we begin, I'll Before we begin, I'll take this opportunity to take this opportunity to take this opportunity to tell you a little tell you a little tell you a little about us. This is our about us. This is our about us. This is our team. We are an team. We are an advanced AI data lab, advanced AI data lab, our team our team our team consists of consists of consists of founders from the founders from the founders from the labs at the labs at the Universities of Washington, Wisconsin, Stanford, and Universities of Washington, Wisconsin, Stanford, and talented in-house talented in-house talented in-house researchers, and we are researchers, and we are researchers, and we are fortunate to work on, fortunate to work on, fortunate to work on, that is, create and that is, create and that is, create and deploy deploy deploy datasets and environments datasets and environments datasets and environments for advanced AI. It is a for advanced AI. It is a for advanced AI. It is a combination of combination of combination of academic and academic and academic and applied applied applied research, research, research, based on based on based on corporate corporate corporate implementations. We implementations. We implementations. We have a team have a team have a team focused on focused on focused on creating the best
-
creating the best creating the best data to develop data to develop data to develop cutting-edge AI. We cutting-edge AI. We cutting-edge AI. We do this in three do this in three do this in three ways. The first is ways. The first is ways. The first is open research open research open research and benchmarks. We are and benchmarks. We are and benchmarks. We are constantly constantly constantly contributing to contributing to contributing to open source testing, and the open source testing, and the open source testing, and the project I will project I will project I will talk about today talk about today talk about today falls into this falls into this falls into this area. I also want to area. I also want to area. I also want to mention the advanced " mention the advanced " live" data, live" data, live" data, environments, and environments, and environments, and enterprise enterprise enterprise implementations implementations implementations we are working on. Within the framework of we are working on. Within the framework of we are working on. Within the framework of open open open research and research and research and benchmarks, we benchmarks, we benchmarks, we support scientific support scientific support scientific activities with grants, the activities with grants, the activities with grants, the possibilities of which I will possibilities of which I will possibilities of which I will talk about at the end. talk about at the end. So get ready. We So get ready. We also engage in also engage in also engage in co-authorship with co-authorship with co-authorship with partners from partners from partners from academia or academia or academia or industry, industry, industry, conduct independent conduct independent conduct independent research, and research, and research, and publish open publish open publish open materials. Today I'll materials. Today I'll materials. Today I'll talk about a talk about a talk about a partnership with the Sky partnership with the Sky partnership with the Sky Computing Lab at the Computing Lab at the University of California, University of California, Berkeley: we created a Berkeley: we created a Berkeley: we created a 4 billion parameter Qwen model 4 billion parameter Qwen model that outperformed the that outperformed the 235 235 235 billion version on billion version on billion version on financial problems financial problems . The model with 4 billion . The model with 4 billion . The model with 4 billion parameters, spoiler alert, parameters, spoiler alert, parameters, spoiler alert, achieved almost 60% achieved almost 60% success in tests, success in tests, success in tests, while the model with 235 while the model with 235 while the model with 235 billion—51%. So, billion—51%. So, billion—51%. So, the results are impressive.
-
the results are impressive. the results are impressive. This was all done This was all done This was all done in collaboration with—and I have to in collaboration with—and I have to in collaboration with—and I have to give credit to— give credit to— Manon Rupta, Manon Rupta, Manon Rupta, Shujian Tan from Sky Shujian Tan from Sky Shujian Tan from Sky Computing Lab, and Computing Lab, and Computing Lab, and Bhavishya, Chris Bhavishya, Chris Bhavishya, Chris Glaze, and myself from Glaze, and myself from Glaze, and myself from Snorkel. You can find Snorkel. You can find Snorkel. You can find all the resources, all the resources, all the resources, training scripts, training scripts, training scripts, and synthetic data on and synthetic data on and synthetic data on our GitHub, so our GitHub, so our GitHub, so check them out. All of this check them out. All of this check them out. All of this is open is open is open source and source and source and available to you. available to you. So, today I’ll So, today I’ll talk about what talk about what talk about what we think we think we think financial institutions financial institutions financial institutions need from AI, need from AI, need from AI, based on our based on our based on our experience with experience with experience with enterprise implementations and the enterprise implementations and the FinQA open benchmark. This is FinQA open benchmark. This is expert-validated data for expert-validated data for expert-validated data for financial questions financial questions financial questions and answers on which and answers on which and answers on which we trained and we trained and we trained and evaluated our evaluated our evaluated our model. And finally, the model. And finally, the model. And finally, the process of training the process of training the process of training the RLLM FinQA 4B model. So let's RLLM FinQA 4B model. So let's RLLM FinQA 4B model. So let's begin. This is the only begin. This is the only begin. This is the only slide with a reference slide with a reference slide with a reference to Charles Dickens, I to Charles Dickens, I to Charles Dickens, I promise, but I would call promise, but I would call promise, but I would call it a tale of two it a tale of two it a tale of two models. First, a models. First, a models. First, a large station wagon large station wagon large station wagon versus a small versus a small versus a small specialist. The bottom line is specialist. The bottom line is specialist. The bottom line is that when that when that when financial institutions financial institutions financial institutions ask us ask us ask us how well a how well a how well a model or agent model or agent model or agent reasons, they usually reasons, they usually reasons, they usually mean, mean, mean, can it work with can it work with can it work with my stack of complex my stack of complex my stack of complex schemas, legacy APIs, and schemas, legacy APIs, and schemas, legacy APIs, and technical debt? And technical debt? And technical debt? And can we trust can we trust can we trust them in realistic and them in realistic and them in realistic and long long long workflows without workflows without workflows without accumulating errors?
-
accumulating errors? accumulating errors? And can we And can we And can we reliably verify the reliably verify the reliably verify the agent's performance? This is agent's performance? This is agent's performance? This is important, for example, in the important, for example, in the important, for example, in the financial and financial and financial and legal fields. This legal fields. This legal fields. This brings us to brings us to brings us to the question: how can we the question: how can we the question: how can we evaluate and evaluate and evaluate and improve improve improve agents to make agents to make agents to make them reliable in these them reliable in these them reliable in these real-world settings? So real-world settings? So , the reflex , the reflex , the reflex reaction may be reaction may be reaction may be scaling. We've all scaling. We've all scaling. We've all seen the results of seen the results of scaling laws, but scaling laws, but today I'll today I'll today I'll argue that argue that argue that specialized specialized specialized processes processes processes don't really need don't really need don't really need universals. We don't universals. We don't universals. We don't need polymaths. You need polymaths. You need polymaths. You know, you wouldn't know, you wouldn't know, you wouldn't call Terence call Terence call Terence Tao for a tax Tao for a tax Tao for a tax audit if you were audit if you were audit if you were lucky enough to have him in your lucky enough to have him in your lucky enough to have him in your contacts. You would contacts. You would contacts. You would call a call a call a specialist who specialist who specialist who knows the forms, knows the forms, knows the forms, tools, and tools, and tools, and rules. That's why we rules. That's why we rules. That's why we created FinQA. This is created FinQA. This is expert-validated data for expert-validated data for financial questions financial questions financial questions and answers. And I'll and answers. And I'll and answers. And I'll walk you through the walk you through the walk you through the multi-stage multi-stage multi-stage pipeline of how we pipeline of how we pipeline of how we built this built this built this benchmark and benchmark and benchmark and training data. training data. training data. So, the first step is to So, the first step is to So, the first step is to extract the schemas and extract the schemas and extract the schemas and data. They come data. They come data. They come from 10-K reports. These are annual from 10-K reports. These are annual from 10-K reports. These are annual reports required by the SEC reports required by the SEC reports required by the SEC for all public for all public for all public companies, and they are companies, and they are companies, and they are used used used by analysts to by analysts to by analysts to identify risks and identify risks and identify risks and make make make investment decisions investment decisions . We take them from the . We take them from the . We take them from the Edgar system and Edgar system and Edgar system and use the use the use the Qwen 3 30B model to Qwen 3 30B model to Qwen 3 30B model to create tables, SQL create tables, SQL tables, about 6,900 tables, about 6,900 tables, about 6,900 of them. And we of them. And we of them. And we use each use each use each table to table to table to create one create one create one question- question- answer pair. This answer pair. This answer pair. This is created together with a is created together with a is created together with a taxonomy of questions taxonomy of questions , developed in conjunction , developed in conjunction , developed in conjunction with financial with financial with financial experts who experts who experts who work with these work with these work with these documents. They are
-
documents. They are documents. They are used as used as used as context for context for context for the model to generate the model to generate the model to generate question- question- answer pairs along with answer pairs along with answer pairs along with metadata, the metadata, the metadata, the purpose of which we will purpose of which we will purpose of which we will learn about later. And learn about later. And learn about later. And finally, the third step finally, the third step is verification. This is a is verification. This is a is verification. This is a three-level three-level three-level verification process. The first is a verification process. The first is a verification process. The first is a programmatic programmatic programmatic consistency check. This is when consistency check. This is when consistency check. This is when we use we use we use metadata, including metadata, including metadata, including origin, origin, origin, table and column names, table and column names, table and column names, to make sure to make sure to make sure we're not making it up. we're not making it up. we're not making it up. The second is The second is The second is automated automated automated checks from checks from checks from independent agents. And independent agents. And finally, expert finally, expert finally, expert manual verification. manual verification. manual verification. That is, people That is, people That is, people review it themselves to review it themselves to review it themselves to make sure the make sure the make sure the questions and questions and questions and answers are answers are answers are realistic and realistic and realistic and verified. And verified. And verified. And the result is the the result is the the result is the Snorkel FinQA dataset. This is our Snorkel FinQA dataset. This is our Snorkel FinQA dataset. This is our first approach to first approach to first approach to this. Realistic this. Realistic this. Realistic queries that require queries that require queries that require planning, planning, tooling, tooling, function calling, and function calling, and function calling, and reasoning. And reasoning. And reasoning. And the answers are the only, the answers are the only, the answers are the only, proven, and proven, and proven, and final results final results . And from this we . And from this we . And from this we get from an get from an get from an initial set of initial set of initial set of about 7000 tables: about 7000 tables: about 7000 tables: 4000 for training, 500 4000 for training, 500 4000 for training, 500 for validation and 290 in the for validation and 290 in the for validation and 290 in the deferred deferred deferred benchmark. The data benchmark. The data benchmark. The data is partitioned so that is partitioned so that is partitioned so that no company in no company in no company in the benchmark the benchmark the benchmark overlaps in overlaps in any of the partitions. Using any of the partitions. Using any of the partitions. Using this data, we found, this data, we found, this data, we found, even in advanced even in advanced even in advanced models, systemic models, systemic models, systemic gaps and typical gaps and typical gaps and typical errors. The first is errors. The first is errors. The first is hallucinations with a
-
hallucinations with a hallucinations with a data schema. The model data schema. The model data schema. The model may assume that may assume that may assume that tables or tables or tables or column names exist, and column names exist, and column names exist, and this may be this may be this may be an artifact of an artifact of an artifact of prior prior prior training, something it has training, something it has training, something it has seen before. seen before. seen before. The next thing is The next thing is context overflow. That is, context overflow. That is, cluttering your own cluttering your own cluttering your own context with poorly context with poorly context with poorly formulated formulated formulated queries. For example, queries. For example, queries. For example, a call to "select star" a call to "select star" a call to "select star" could exhaust the could exhaust the could exhaust the context limits context limits context limits if successful. And if successful. And if successful. And finally, the inability finally, the inability finally, the inability to recover from to recover from to recover from mistakes. When a model mistakes. When a model mistakes. When a model repeats the same repeats the same repeats the same failed strategy failed strategy failed strategy instead of instead of instead of analyzing analyzing error messages and error messages and adapting. These types of adapting. These types of adapting. These types of errors errors errors are seen not are seen not are seen not only in the financial only in the financial only in the financial sector, which I am sector, which I am sector, which I am talking about today, talking about today, talking about today, but also in insurance and but also in insurance and but also in insurance and underwriting. This is the underwriting. This is the underwriting. This is the work I want to work I want to work I want to highlight, it was highlight, it was highlight, it was presented at the presented at the presented at the CAIS conference CAIS conference CAIS conference a few weeks ago. a few weeks ago. So, now I'll move on So, now I'll move on to training the R to training the R to training the R LLM FinQA model on 4 billion LLM FinQA model on 4 billion LLM FinQA model on 4 billion parameters. To do this, parameters. To do this, parameters. To do this, we used the we used the we used the R LLM framework, R LLM framework, R LLM framework, developed developed developed by collaborators from the by collaborators from the by collaborators from the Sky Computing lab.
-
Sky Computing lab. If you haven't If you haven't used it yet, it's a used it yet, it's a used it yet, it's a very handy very handy very handy tool, and I want to tool, and I want to tool, and I want to highlight a few highlight a few highlight a few key key key features: it features: it features: it works with any works with any works with any agent framework agent framework , requires minimal , requires minimal , requires minimal code changes thanks to the code changes thanks to the code changes thanks to the decorator pattern, decorator pattern, decorator pattern, has a CLI-oriented has a CLI-oriented has a CLI-oriented workflow, and workflow, and workflow, and proven proven proven results. We have seen results. We have seen how R LLM improves how R LLM improves how R LLM improves results not only in results not only in results not only in this financial this financial this financial case, but also, for example, case, but also, for example, case, but also, for example, in mathematical in mathematical in mathematical reasoning. reasoning. Several RL algorithms are built in Several RL algorithms are built in , such as GRPO , such as GRPO , Reinforce RL, OOO, which can be , Reinforce RL, OOO, which can be , Reinforce RL, OOO, which can be customized, and customized, and various various training backends are supported. Okay, training backends are supported. Okay, training backends are supported. Okay, now about our now about our now about our learning environment. learning environment. learning environment. Um, this is, uh, the Um, this is, uh, the Um, this is, uh, the learning environment that we learning environment that we learning environment that we used for used for used for our model. The agent our model. The agent our model. The agent runs in the react loop and runs in the react loop and runs in the react loop and has access to has access to has access to tools to tools to tools to interact with interact with interact with the tables we the tables we the tables we generated. generated. generated. The environment includes The environment includes The environment includes about 7000 about 7000 about 7000 tables that we created, and tables that we created, and tables that we created, and the reward is the reward is the reward is binary correctness binary correctness , which is determined by a language , which is determined by a language , which is determined by a language model judge, in this model judge, in this model judge, in this case GPT-5 nano, using a case GPT-5 nano, using a benchmark-based score.
-
Training details Training details are provided here. This, uh, are provided here. This, uh, are provided here. This, uh, will probably be interesting to will probably be interesting to will probably be interesting to people. We have a basic people. We have a basic people. We have a basic Qwen 3 4B model, GRPO and Qwen 3 4B model, GRPO and Qwen 3 4B model, GRPO and binary reward binary reward binary reward for optimization and RL for optimization and RL framework. We framework. We framework. We ran approximately ran approximately ran approximately 1,000 parallel 1,000 parallel trajectories , and it all cost , and it all cost , and it all cost less than $500 to less than $500 to less than $500 to get the get the get the results we results we results we saw. Breakdown: saw. Breakdown: saw. Breakdown: approximately $420 approximately $420 approximately $420 for computation and for computation and for computation and $40 per judge, $40 per judge, $40 per judge, i.e. per judge API. And that i.e. per judge API. And that i.e. per judge API. And that was on eight H100s. was on eight H100s. The first result is the The first result is the central question: central question: central question: can a 4 billion can a 4 billion can a 4 billion model, equipped with model, equipped with model, equipped with special special special tools and tools and tools and training, truly training, truly training, truly compete with its compete with its compete with its larger counterpart? And larger counterpart? And larger counterpart? And the answer is yes. So, the answer is yes. So, the answer is yes. So, we saw that we saw that we saw that the accuracy of the base the accuracy of the base the accuracy of the base model more than model more than model more than doubled, and it doubled, and it doubled, and it outperformed its outperformed its outperformed its larger version, the larger version, the larger version, the 235 235 235 billion parameter model billion parameter model . And this despite the fact that . And this despite the fact that . And this despite the fact that it is only a fraction of it is only a fraction of it is only a fraction of its size. And we its size. And we its size. And we evaluate all the evaluate all the evaluate all the models shown here models shown here models shown here on a deferred on a deferred on a deferred set of 290 expertly set of 290 expertly set of 290 expertly selected samples.
-
selected samples. So a natural So a natural concern is whether concern is whether concern is whether this extends to more this extends to more this extends to more complex tasks. complex tasks. complex tasks. After all, these are just pairs of " After all, these are just pairs of " one question—one one question—one one question—one answer" for answer" for answer" for one table. To one table. To one table. To investigate this, we investigate this, we investigate this, we examined our data examined our data examined our data for for for FinQA inference. This is an FinQA inference. This is an FinQA inference. This is an evolution of the evolution of the evolution of the FinQA dataset that I FinQA dataset that I FinQA dataset that I presented earlier. presented earlier. The difference here is that The difference here is that now we have now we have now we have dependencies on dependencies on dependencies on many tables, from many tables, from many tables, from two to five, and this two to five, and this two to five, and this requires consistent requires consistent requires consistent decision-making and decision-making and decision-making and planning. And we planning. And we planning. And we conducted an evaluation on conducted an evaluation on conducted an evaluation on this dataset. this dataset. this dataset. You can see You can see You can see the distribution as well. We have the distribution as well. We have the distribution as well. We have about 1000 for about 1000 for about 1000 for training, 120 for training, 120 for training, 120 for validation, 80 in validation, 80 in validation, 80 in benchmarks. And we benchmarks. And we benchmarks. And we see that the discipline of see that the discipline of see that the discipline of using using using tools tools has actually generalized directly has actually generalized directly from learning on simple from learning on simple from learning on simple data. And no data. And no data. And no special special special training on training on training on multi- multi- multi- table examples was really table examples was really table examples was really needed to needed to needed to see the gains in see the gains in see the gains in this this this multi-table multi-table multi-table variant of the variant of the variant of the dataset. We will investigate dataset. We will investigate dataset. We will investigate why this is so why this is so why this is so using ablation using ablation using ablation studies.
-
studies. Our next Our next concern was concern was concern was whether this result would extend whether this result would extend whether this result would extend to more to more to more general general general use cases. To do use cases. To do use cases. To do this, we used the this, we used the this, we used the BFCL benchmark, which BFCL benchmark, which BFCL benchmark, which measures the overall measures the overall measures the overall capabilities of capabilities of capabilities of tool calls. And the tool calls. And the tool calls. And the overall accuracy overall accuracy overall accuracy actually improved a little bit actually improved a little bit actually improved a little bit , and then , and then ... We also got ... We also got ... We also got some small gains in some small gains in some small gains in multi-turn multi-turn multi-turn dialogues and memory. dialogues and memory. dialogues and memory. Therefore, specialization Therefore, specialization Therefore, specialization did not impair the did not impair the did not impair the model's broader competence in model's broader competence in model's broader competence in using using using tools when we tools when we tools when we performed RL performed RL tuning for tuning for tuning for this model. So, to this model. So, to this model. So, to understand what understand what understand what drove this drove this drove this performance, we performance, we performance, we conducted conducted conducted ablation studies on ablation studies on ablation studies on different different different data combinations, i.e., ways to data combinations, i.e., ways to data combinations, i.e., ways to combine the available combine the available combine the available training data. training data. training data. The first is training The first is training The first is training on FinQA data only, the on FinQA data only, the on FinQA data only, the second is on tables ( second is on tables ( single and multiple) single and multiple) , and finally, a , and finally, a training program training program training program where we start with where we start with where we start with one table and one table and one table and then move on to then move on to then move on to multiple. And, surprisingly, it was multiple. And, surprisingly, it was multiple. And, surprisingly, it was training training training on a simpler on a simpler on a simpler dataset that yielded the dataset that yielded the dataset that yielded the biggest gains. biggest gains. This partly explains the generalization results we generalization results we saw earlier. The saw earlier. The saw earlier. The bottleneck was bottleneck was bottleneck was not the depth of reasoning, not the depth of reasoning, not the depth of reasoning, but the use of but the use of but the use of tools, that is, this is where tools, that is, this is where tools, that is, this is where the model the model the model failed, and failed, and failed, and we worked to we worked to we worked to improve improve improve the reliability of this the reliability of this the reliability of this aspect. And once you aspect. And once you aspect. And once you master the basics, you master the basics, you master the basics, you can combine these can combine these can combine these skills. Yes, oh, skills. Yes, oh, skills. Yes, oh, let me let me let me go back. The go back. The next logical next logical question is whether question is whether question is whether more informative more informative more informative reward signals, reward signals, reward signals, such as such as such as encouraging intermediate encouraging intermediate encouraging intermediate steps,
-
steps, steps, table access, table access, table access, query completeness, etc., can query completeness, etc., can query completeness, etc., can accelerate learning. accelerate learning. accelerate learning. This, for example, is This, for example, is This, for example, is different from the different from the different from the curriculum. curriculum. curriculum. The most complex The most complex The most complex option we option we option we developed is a kind of developed is a kind of developed is a kind of rubricator. It rubricator. It rubricator. It uses a uses a uses a detailed detailed detailed assessment with assessment with assessment with several weighted several weighted several weighted components. By the way components. By the way , we developed it , we developed it , we developed it together with experts. together with experts. together with experts. And despite such And despite such And despite such investments in investments in investments in reward engineering, reward engineering, reward engineering, we have seen again that we have seen again that we have seen again that simpler is better. simpler is better. It was the single binary It was the single binary reward signal that reward signal that reward signal that provided us with the provided us with the provided us with the best result best result best result on this data. This on this data. This on this data. This led us to the led us to the led us to the concept we are concept we are concept we are developing for developing for developing for corporate corporate corporate agents. Overall, this agents. Overall, this agents. Overall, this study showed study showed that small models that small models that small models can be successfully can be successfully can be successfully trained trained trained using RL if using RL if using RL if there is quality data to there is quality data to there is quality data to represent the represent the represent the target outcomes target outcomes target outcomes and tasks. This can and tasks. This can and tasks. This can lead to better lead to better deployment economics for deployment economics for specialized specialized specialized tasks and domains. We have tasks and domains. We have tasks and domains. We have observed and observed and observed and confirmed this confirmed this confirmed this result in other result in other result in other areas, including areas, including areas, including healthcare, healthcare, healthcare, law, and insurance, law, and insurance, law, and insurance, which I pointed out which I pointed out which I pointed out earlier. So earlier. So earlier. So get to know them.
-
get to know them. Our point of view on the Our point of view on the formation of criteria formation of criteria formation of criteria for evaluating for evaluating for evaluating agents is as agents is as agents is as follows. follows. The complexity of The complexity of the environment is one the environment is one the environment is one of the most important of the most important of the most important criteria. How criteria. How criteria. How complex, complex, complex, realistic, and realistic, and realistic, and dynamic are the dynamic are the dynamic are the environments in which environments in which environments in which agents operate. agents operate. agents operate. The next criterion is the The next criterion is the horizon horizon horizon of autonomy. Are of autonomy. Are agents tested and evaluated at agents tested and evaluated at the levels and the levels and the levels and durations that durations that users need to users need to work with them? Do work with them? Do work with them? Do they make they make they make safe and safe and safe and sound decisions on sound decisions on sound decisions on their own? And their own? And their own? And can they can they can they improve their improve their improve their judgment as a “co- judgment as a “co- judgment as a “co- pilot” over time? And lastly, the pilot” over time? And lastly, the complexity of the source complexity of the source complexity of the source data. We want data. We want data. We want to cover a wide to cover a wide to cover a wide range of results from range of results from range of results from our daily our daily our daily work. Hmm, I would say work. Hmm, I would say work. Hmm, I would say this area is currently this area is currently this area is currently under- under- under- researched. researched. New New benchmarks are emerging in this benchmarks are emerging in this benchmarks are emerging in this direction, but I still direction, but I still direction, but I still see the main focus see the main focus see the main focus on text on text on text artifacts, and artifacts, and artifacts, and creating more creating more creating more rigorous assessments in these rigorous assessments in these rigorous assessments in these conditions is a big conditions is a big conditions is a big opportunity. There opportunity. There opportunity. There is is is progress here. So, with that in progress here. So, with that in progress here. So, with that in mind, I want to mind, I want to mind, I want to announce our announce our announce our open benchmark grants that open benchmark grants that funded this funded this project. This is a project. This is a project. This is a multi-million dollar multi-million dollar multi-million dollar investment in investment in investment in creating such creating such creating such open benchmarks open benchmarks . We have made significant . We have made significant . We have made significant progress with Agent, Last Exam, progress with Agent, Last Exam, progress with Agent, Last Exam, Judgment Bench, and others Judgment Bench, and others Judgment Bench, and others emerging now, and emerging now, and emerging now, and this greatly accelerates this greatly accelerates this greatly accelerates our ability to
-
our ability to our ability to measure and measure and measure and shape cutting-edge shape cutting-edge shape cutting-edge technology. So, to technology. So, to technology. So, to apply, apply, apply, visit our visit our visit our website. You will go website. You will go website. You will go through the process of submission, through the process of submission, through the process of submission, selection and selection and selection and review, review, review, development, and ultimately development, and ultimately publication, launch, and publication, launch, and publication, launch, and promotion. I also promotion. I also promotion. I also want to note that we want to note that we want to note that we hire professional hire professional hire professional researchers, researchers, researchers, engineers, and other engineers, and other engineers, and other specialists, so please specialists, so please specialists, so please pay attention to this. pay attention to this. You can You can browse our website, browse our website, browse our website, and in conclusion— and in conclusion— and in conclusion— thank you very much for participating. thank you very much for participating. You can You can use this QR code use this QR code to see the blog, to see the blog, to see the blog, as well as links to as well as links to as well as links to all the artifacts all the artifacts all the artifacts published on GitHub published on GitHub published on GitHub and Hugging Face. Yes, thank you and Hugging Face. Yes, thank you and Hugging Face. Yes, thank you all.
No summary available yet.
View original episode ↗