AWS’s Ankur Mehrotra on Building Generative AI Models

Transcription

hi I'm James Maguire and on today's e speaks we're talking about generative Ai and building artificial intelligence models we're taking a look at what companies need to know about these emerging Technologies to gain the most competitive advantage to discuss that I'm joined by a major industry expert with me is Encore motra general manager of sage maker at AWS Encore actually thrilled to have you with us today thanks James it's great to talk to you you know I'm amazed at the amount of interest in generative AI obviously it's it's exploded under the scene what are the trends driving the market for generative AI here in mid 2024 in your conversations with customers and I know you do talk with a lot of customers what matters most of companies and and what are companies looking for in terms of a foundational AI model yeah James uh as you know generative AI has really been a transformational uh uh technology you our in our industry and uh we're seeing customers move quickly to apply generative VI uh to their existing user experiences to improve us experiences but we think that uh you know there going to be new user experiences that will also be enabled with the use of this technology and uh what we're seeing is that this the the pace of innovation here is um is is amazing and there are more and more uh powerful models being built um in every few weeks now and uh what we hear from customers is that they really want to use uh experiment and apply the best uh models uh to the use cases so customer customers are looking for tools to be able to experiment and uh test these models and then move them to uh to production use quickly another thing that we're hearing from customers is that uh they want to use these models but they want to do it in a way uh that ensures security and uh preserves their data privacy and that's something that they have trusted AWS for since many years so many um customers are looking to use generated models on AWS well I guess it's it's not always there's certainly going to be a learning curve here so I guess i' ask you what are some of the key challenges the customers are facing as they build these stand of AI applications and one of the things you talked about getting into production is is that an issue and Andor you know privacy and compliance it seems like that would be an issue that that companies grapple with but what do you see the the real headaches yeah there there are a number of challenges that customers uh face while applying this technology to the use cases so the first one is that uh there are there are many different models that are available now and uh and what we're seeing now and this is something that that AWS was the first to to highlight is that no one model is actually sufficient to solve or address all use cases so customers are finding that one model may be better at building a particular uh kind of user experiences let's say a chat based application versus another generative AI model may be better at uh assisting with coding or or software development so uh our customers are finding that they need to use not just one model but a combination of different models and they're looking for um tools that can help them move test and move these models to production use quickly another thing that uh our customers are a challenge that they're facing is that uh understanding the risks with the use of these models is is important so you may have heard about hallucinations uh as as a problem with Genera VI models where sometimes they can you know make things up and and there are various other risks that are customers want to understand before they put these models to use for a serious or a mission critical use case so for instance they may want to understand um you know the the Model Behavior along a different set of responsibil parameters such as bias or toxicity or sem semantic robustness so uh being able to comprehensive ly evaluate a new generative AI model and then being able to control for those risks while applying the model in production is is a is a key challenge at the same time we're seeing that many times customers want to adapt these models to their own domain and which requires taking their own data and then fine-tuning the model or using other AI techniques such as in context learning to be able to tweak the model and have the model adapted to their uh unique use case or the or domain and then using the model um for for their actual use case and being able to do that uh can be you know timec consuming today and here at AWS we are building new tools and experiences to help our customers go through that Journey quickly and easily and lastly James as you mentioned security and privacy is really important so we are uh focused on help our customers use these models in a secure and private way so that their data is uh protected at all times and they're able to safely deploy these models in production and integrate them uh with their applications it interesting I I think it seems like I wonder if one of the challenges is the companies in some cases May lack the expertise inhouse they they want to build the model but they need a helping hand do you run into a lot of that in terms of customers they they don't have expertise in house many times so we uh do you know we are really focused on helping customers with all skill sets to be able to apply this technology uh to their business problems and we are seeing that uh you know customers I talked about model evaluation as a as a capability and customers often times don't have the uh the ability to to implement algorithms that can comprehensively evaluate these models for production use or understand the risks which is why um we we launched capabilities for um you know model evaluation through uh Amazon sagemaker clarify a few months ago at our reinvent conference through which customers can take a generative AI model and run it through this tool that we launched and uh see a you know a comprehensive report around how the model performs across different uh responsible lay parameters mhm and we also see that uh as I mentioned customers want to bring their own data to you know fine-tune these models and sometimes customers don't know where to where to start and how to get started so we provide simple apis for customers to get started pick a model and um and and be able to tweak certain parameters and then adapt the model to their use case and then finally uh easily deploy it uh deploy the model also to their us use case in a way where the model can scale the the model inference as it it is referred to can easily scale with the uh growing you know demand for the end user application that they're building or applying the model to so all these are tasks that are involved in the AI life cycle which customers are looking to to simplify and automate to the extent possible and that's uh really what we're focused on um making um making it easier for our customers M I thought I heard at one point you talked about in essence a pre-fabricated model pre-built model that companies could customize so they're taking that pre- pre-built model and simply putting in their own data is that generally how that would work or including their own parameters so we're seeing uh you know this space is really and the technology is evolving it's still very early and so we're seeing customers apply a few different techniques to adapt the models or augment them with their own data so uh some of our customers they have um you know unique data sets which are uh which you know they have acquired through uh you know unique interactions between their customers and and and their applications um oftentimes the do you know data is domain specific let's say for a particular particular vertical let's say Financial Services or Healthcare Etc and so customers are able to take this data and then uh and go through a process called fine-tuning which is essentially take about taking a an already pre-trained model and then training it further with their proprietary data to adapt the model to their own domain so we're seeing customers do that so uh sagemaker provides uh apis for customers to be able to do this um easily and and securely and uh so we seeing customers you know appli uh use fine tuning as as a technique some customers are also using techniques uh such as rag which is retrieval augmented generation right it's technique where you know you can these models accept some have something called as context window so these through which the models can accept context in which they're expected to generate the response so our customers are also using that as a technique rather than fine-tuning to also adapt these models to the data so we really seeing you know a number of different approaches being tried and experimented with and uh here at AWS we're really focused on giving our customers choice and freedom to choose different approaches and giving them the best tool for the for the right job well on that very subject let's drill down into the S maker offering I know that that plenty of leading generative AI foundational models are trained on Amazon sagemaker so what what sets the solution apart I'll give you a chance to to brag about your solution a little bit thank you James yeah so sagemaker is now you know we we launched the service uh more than six years ago and uh it is now used by hundreds of thousands of AWS customers um to build train deploy uh AI models including generative AI models and uh it is the most um comprehensive service in terms of capabilities that are purpose built for every step in the life cycle so whether it's preparing data to to build a model or uh creating a model and then training it which requires you know specific set of tools uh to do that at scale and then taking that model and then deploying it in a cost efficient way and in a way where the model usage or inference can really scale with the incoming demand and then also after that uh being able to manage a model that has that is already being used in production so for example uh monitoring whether a model is continuing to act or behave in a way it's supposed to and being able to detect if anything in production is changing compared to what is expected um so those are referred to as AI Ops processes and so sagemaker has purpose-built tools for every step in of um of of of the AI life cycle and for every Persona that is involved you know as a few years ago um AI was uh was mostly a data scientist activity where there was a data scientist who was building these models and over the years the number of uh personas that are involved in building end to-end aias Solutions has really increased so uh from data scientists we had you know ml Engineers um get involved who were became responsible for taking these mod and then deploying them into production and then we had um other St business stakeholders got involved to help convert a business problem into an ml problem and then we had data engineers get involved to help prepare the data so at with sage maker we've really focused on working backwards from our customers understanding the need and building the right tool for the for the right job and for the right Persona I think you've talked about this a little bit but I want to make sure I get the full story how is sagemaker helping organiz ganizations build transformative gen AI Solutions easier said than done especially considering as you said the market is very new just want to make sure I get the full story on that yeah James so we're seeing customers across various domains and sizes uh use sagemaker to do generative AI at scale and build generative AI based Solutions so for example if you take uh if you consider uh model training uh some of the you know B best in state-of-the-art generative AI models have been trained using sagemaker so customers like uh perplexity AI or um hugging face or AI 21 Labs uh LG AI research they've built their generative AI models using sagemaker and what they like about Sage maker is the fact that it gives you a managed uh set of capabilities that can scale automatically to help train these model so one of the challenge with training these generative VI models is that they require a large amount of compute so uh there's a model training task and it gets distributed across a large number of um you know compute instances that have these accelerated you know compute uh Al you know Common ones being Nvidia gpus for instance or ews uh own silicon such as um ad tranium and um and how these models get trained across a large amount of compute um requires you know specialized skills and it takes you know time and resources to do that job so uh purpose build capabilities and stagemaker help customers achieve that in 40 up to 40% less time so perplexity for example uh reduce the time to train um their generative AI models uh by up to 40% with sage maker and uh we've seen many other customers cers uh train their models through sagemaker as well and then coming to model inference that is where being able to deploy a model and host it for inference in a cost performant way is extremely important and there are various other factors that are important to customers there including uh the latency of models because it inferences how uh you know users actually experience the these models so uh the performance of of of inference in model hosting is really important so uh our customers such as Salesforce are using Sage maker to deploy uh and host generative AI models for using their applications and are achieving up to uh 50% lower latency and then outside of these capabilities if you look at um the need for automation where which are typically known as AI Ops tools where you have uh capabilities such as sagemaker Pipelines and experiments which can really take help you build automated workflows to create endtoend AI Solutions um those are also helping our customers such as NatWest reduce time it takes to go from a use case you may have in mind to go uh to go to production by up to 75% so we've got customers who are using sagemaker to uh to build AI Solutions across you know the whole spectrum and uh and there achieving significant um improvements in in time and cost and performance well I think the the big question is where are we going in the future of this companies really want to know I mean clearly this sector is evolving at a blistering rate what do you see as the near to midterm future of generative AI James so what we're seeing is that there are first of all the pace of innovation here is going to continue and we're going to see a lot of different types of models being built some of the you know we hear that these generative AI models are getting just getting larger and more powerful but that is true but we're also seeing other smaller Tas specific models also being created so our customers you know when I talk to customers what I hear is that they are having to now they foresee having to use multiple models uh some maybe task specific and some maybe you know others that are more more generalized together to achieve um you achieve their goals and and the ability to do that quickly and safely and securely is is very important to them the second thing that we are noticing is that um models are Mo models are turning into model systems so our some of our customers are looking into using a collection of models that work in tandem so for example one of our customer is using is deploying um a set of different models where one model is responsible for redacting pii um from uh from text and then another model is taking that text and then summarizing it so uh we're going to see that customers will you know um will will now think of these systems as model systems and use a combination of different models that are to together we are also seeing uh Trends where uh customers want the data to be collocated with these models and really create these model systems that they're using in production um we also think that you know customers are going to need uh tools to manage the tradeoffs that are um involved in use of these models so for example uh for building a highly interactive chat-based uh application latency or infant latency is is really important so customers may choose a model that provides them really you know low latency even if it that comes with a slightly higher cost whereas for a a backend business process you know some customers are okay with a higher latency if they can save costs so I think the tools that can help our customers manage these tradeoffs um for the use cases and applications are going to be really important and that's a key area here uh at AWS Spirit Focus you know just one more followup on that I I think the idea of the multiple models is fascinating I had not heard that one so it means more than one model would be F you know uh flowing into an application would be feeding an application is that um I mean I guess we're going to look back at today where some some applications usually only have like one model I think that's really quite primitive like perhaps in the future most scenarios would have multiple models or actually model systems as opposed to Simply a model model is that is that correct in your view of the future I think so I think certain models will be uh good at one task versus other models will be good at uh a different task and uh there may be you know differences in terms of uh cost and performance and uh relative to the task that the model is being used for so I think this this uh this early Trend that we're seeing of multiple models being used in combination is is here to stay very interesting Encore thank you so much for sharing your expertise today I really learned a lot it's fascinating and uh I hope you will come back and talk with us again sometime thanks James it was great talking te

This transcript was generated automatically from the video's captions and may contain errors.

Verfasst von
James Maguire
James Maguire
Published: Jun 20, 2024
Updated: Sep 23, 2024
1 minute read
eWeek Inhalte und Produktempfehlungen sind redaktionell unabhängig. Wir können Geld verdienen, wenn Sie auf Links zu unseren Partnern klicken. Mehr erfahren

Watch my extended interview with Ankur Mehrotra, GM of SageMaker at AWS, to hear his thoughts about how companies are strategizing to build better AI models, along with a range of other AI-related topics.

James Maguire

James Maguire has been reporting on emerging technology for more than 15 years. He has won two ASBPE Awards of Excellence for in-depth feature articles about cloud computing and artificial intelligence. He has covered the gamut of enterprise and consumer technology, and regularly communicates with leading IT newsmakers, vendors and analysts.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Eigentum von TechnologyAdvice. © 2026 TechnologyAdvice. Alle Rechte vorbehalten

Werbetreibenden-Offenlegung: Einige der auf dieser Website erscheinenden Produkte stammen von Unternehmen, von denen TechnologyAdvice eine Vergütung erhält. Diese Vergütung kann beeinflussen, wie und wo Produkte auf dieser Website erscheinen, einschließlich beispielsweise der Reihenfolge, in der sie erscheinen. TechnologyAdvice schließt nicht alle Unternehmen oder alle auf dem Marktplatz verfügbaren Produkttypen ein.