Welcome to Real Talk on AI and I care with your host, Dr. Scott Morris. And I'm co host, Dr. Rayhan Ahmed.
In each episode, we discuss current news, adding a little AI education, and debate innovative AI tools and topics that will change the industry. Hello, everybody. This is Scott Morris and Rayhan. We are here for a real talk.
You imagine this is real talk 26, Rayhan. It's just the craziest thing that we've got that many in. Hopefully all of you have listened or have listened to a few of these. And that's why you keep coming back.
So we can all learn from each other and have some conversations that maybe aren't always being had in the industry. Yeah, fantastic. I can't believe it's been 26 episodes. I know I'm learning a lot.
And I get notes from many of our listeners. And so happy that we're providing a, I think, a timely service. So, Scott, so tell us what's what's been going on every week on thing and there's can't be none of their new news item. I mean, there's just too big items.
I mean, and maybe more than two, but we're going to talk about two. I think the big announcement today, or this week, I should say, was top con has signed an acquisition agreement for Toku. And I know we've talked about Toku a couple of different times in the field of akeolomics, you know. And what Toku's big thing is is that they are, they're a New Zealand company.
And they've kind of been working on these different platforms looking at how retinal images will build a help detect cardiovascular risk and what are biological ages versus our chronological age and looking at kidney disease. Once again, we keep coming back to that same story over and over again, right on its oculomics is maybe the word of this of this year is it's a huge advancement and top con is saying, hey, I'm going to acquire Toku and get our feet wet in this far, this area of oculomics. And they've already got a couple other acquisitions and are different or feet and other things, but I think it's interesting to see some of the big players looking at the future is not being I care is just I care, but I care in terms of, hey, we may be the gatekeepers to many forms of systemic disease and maybe even looking at from the form of preventative disease as we may build to see things and go, hey, you're showing really some of the things that we're looking at. You're showing early signs, let's talk about prevention before we really get to systemic macro, you know, disease states, maybe we can still catch it at microscopic levels, so I think that that's, I mean, to me, that's pretty interesting and then they also introduced and I guess, you know, you talked about earlier tonight when we're in our prep meeting is that, you know, is that this is maybe not a new news, but maybe a reintroduction is top con back in the news again on this one is kind of introduced.
So, I'm just saying IDHA, which is kind of there, oh, let's call it their institute of digital health, right, so they're trying to say we want to be the clinical research center and build a data ecosystem for looking at disease, not just an eye care, but across medicine. So, I think that's fascinating that hey, we're kind of getting into not only a colomics as we've just talked about, but really somebody focusing on the fact that we need a way to build a house and look at data and we're going to talk about that later tonight in our fireside of what that really means. But I think you look at those two big announcements from or announcement and re announcement from top con and I start to see you know if we start looking at where the big players are going, we start seeing a picture that's much more than just in the lane I care. Yeah, after super fascinating, you know, top, top con it's a fascinating company recently, you know, acquired or in talks with it was obviously the KKR.
And I have a top consul lamp and for many, many years, the top con was kind of like a fund this camera company, slot lamp company, I'm sure a lot of our listeners use their devices. It really seems like it's, I don't want to say morphing in, but becoming much more of a software or SaaS based type of. Well, I think you and I talked about it last week right is your favorite phrase is data is the new oil. I think they're starting to their where ones companies are going maybe we need to be worrying more about the data that we're acquiring from through software and through diagnostic technology and then best being a technology hardware vendor.
Yeah, super excited for for toku. I think there are, you know, their success is I think AI and I care is success. Of course, I feel only makes it a deep foundation for that I do just to plug our journal AI and I care.com I recently wrote on some of the challenge of the box. I'm a big believer in it, but they're definitely challenges.
So to check out what some of those might be please check out our recent my recent column there on AI and I care. I'm going to talk about AI and I know Scott in our pre meeting you were you've been talking about this to this is something called retrieval rags or retrieval augmented generation. And I mentioned really like I love doing these podcasts because it's a way for me to really learn AI. I'm not an expert by any stretch, but I love using these tools and they're advancing so fast.
And we're all I'm even the experts are learning as we're going. So this is, you know, we're getting more and more, I think advanced in these knowledge by Scott as we're starting building layer by layer. We started with what is AI and neural networks and now we're talking about pretty sophisticated topics like rack. So.
And I think last but great point is it's layering and we're starting to see a lot of companies layer rags on top of other existing AI. So you're going to be coming multi modal AI, if you will. So I'm super is this topic. I'm super interested in around teach me.
Yeah. So, you know, so in last week we talked about hallucinations, right. This is and we've all experienced it before. I mean I explained it.
I think daily every other like when I'm using chat GPT, which I you know. I have that I have that premium package that I use it so much. I mean they could sound completely confident and so entirely wrong and it happened because they don't have the access to facts. And real time, right.
Like they just sort of they want they have alignment problem. They were they want to please their master and they generate answers from patterns they've learned during training and training that ends at a fixed point time. So this brings us to the concept that's really shaping these next gen AI systems, especially I think in healthcare or domain specific areas where you have access to data and you may not want to share that because I patient protected or regulatory protected. So that's the way that everyone else and so what do you do use something called use tools like rats or a rag which stands for retrieval augmented generation that's retrieval augmented generation.
So let's break that down. So the word generation and that refers to what large language models already do right. We know and they do it well. They take a profit in a generate text and new text but based on learned patterns.
So the retrieval means the system for searches relevant information from a source that's connected and they call them you know so it's going to be like a knowledge base right maybe a research paper guidelines HR data your own documents like your own clint your own clinic documents. So in a rag pipeline when you ask a question you direct the system to first retrieve it from the most relevant content that to generate an answer right so in your prompt and I use rags comment you know very frequently I say please use so and so knowledge base in that order right. Why is this necessary right like well you know why do we need a rag right so because building these foundation models right the underlying engine behind GPT Gemini Claude these are and so we've talked about this extraordinarily expensive and it's getting more expensive to do it. So we're trained on trillions of words tens of thousands of GPUs running for months and we did talk about deep seek many months ago but you know that that chatter has come down a little bit there maybe cheaper weights but it's still super super expensive not every company cannot be doing this right and so you can't just update your new foundation model every time there's new evidence or you want to change something you have to of course build on the back end of these large system and that's the basis for opening eyes value.
So instead of retrained entire model you pull in new information dynamic from your own knowledge base or whatever external data source that you prioritize you don't have to teach it everything again you don't teach it how to like talk or you know how to put it sentences together that's already been done in the foundation model so okay for example right so a clinician might ask you know what's recommended dose for such and such like a citizen's all on my I didn't give me disease or I'm not going to do that. But I can do using right this side that I get into the place or is being followed by the patients microwave patients or something and the rag would then search and each hospital and can hundred may have their own specific answer right another thing may be You know, what does this patient need a pre-authorization based off their insurance and our own policies before we do it in critical, you know, injection, right? So the retriever then looks at the documents synthesize with the model and it gives you an academic summary and plain English, right? So it's not about making the model as a bigger smarter.
It's making them more relevant, verifiable, right? And something that you could actually use in a cost-efficient way, you're not rebuilding obviously JTPT, but you're able to sort of deal with this issue of hallucination, improper information, because you're prioritizing their answers, right? And they're companies that are actually built on, like, open evidence. I was just doing some research.
There you go. That was an example I was going to get. Yeah, so they're being, I mean, they, they, so they don't actually publicly say that they're basically a rag, but they basically are, I mean, they're not, they have not rebuilt JTPT or Gemini. We don't know what back end they're using, it doesn't matter, but they are using new and journal medicine and Mayo Clinic and other sort of well well validated, extra information that the secret sauce for them is they got that permission, right?
So they got all those articles and, you know, no one else can legitimately really like claim to those as far as I know. And so that's where they're, that's where they do. But, you know, any company or clinic and theoretically do this if you have your own internal sets of documents, you can build on top of that. So I hope that was helpful for the audience.
This is something that, you know, I'm learning about every day. And I love this idea of having a more customized LLM that's built purpose specific for me in my use cases. And I think it's going to be something I think ragged for the audience. I think you're going to hear that in more and more and more people as they come out and saying, we're built on this January, we're built on the system.
We lay a rags on the top of it. To make sure that the date is real and it's not being generated. It's fictitious. And I think that's like, well, that's, this is a term that I'm so glad that Rayon went over because I think it's something that we want to hear.
We should hear. I think it's going to be the new standard is if you're not using a rag, then where you really get your data from, right? So. And I bring this to our fireside and you know, Rayon and I, I've been thinking a lot about this this week and I keep I'm this pro AI, I think, as you can get.
But I'm also a realist and I, I look at them people go, yeah, yes, got you talk about this, but there's got to be problems and I look at and go, I think there's two huge, massive problems. We're going to call them the right and left to kill these heal of AI. And I think it too we figure this out AI will probably not get a lot more dramatic and change the world much more that already is. These two problems and they're very distinct and different problems.
And so the first problem we're going to kind of relay on is what. Rayon just talked about his data and kind of what we talked about in, you know, or news of the day, I think that the smart companies are starting to look at this and going, you know, a really big problem is data because right now, especially in healthcare and we'll boil us down to eye care. We have no centralized database of real data. We have studies and, you know, journal articles and stuff like that, but those are end values of hundreds or thousands of people.
And if you think about it in the United States alone, we're seeing hundreds of thousands of examinations every single day. And we have no way to really access that real data and people go, well, yeah, you can scrub any HR, but each ours don't collect all the data. They just collect what's in their check boxes, right? And so that's a long way from what's really happening in terms of data.
And you know, I, I don't know about you guys as listeners, but I always think do I plug in? I always think do I plug every single question I ask and every medical decision I make in my head, do I plug it into my EHR? I'd love to say the answers, yes, but the answers know. And so if we're not writing it down, but we're scrubbing the EHR for the data, is it really an accurate valid representation of what's going on?
The answer is not even close. That's number one. And then number two, as we've talked about this another me other podcast we've done, Rayon is that, you know, then we get in the problem is, can you really scrub the data? Is it is it your data as the HR company to scrub?
And there's some, there's some really big patient confidentiality patient privacy issues there, not just saying, well, it's all anonymized, but did they give you permission to share that data? And the answer is they didn't. There's going to be some really big legal battles over companies taking data from EHRs without patient permission or patient transparency in terms of how that data is being used. And I think that that's a whole nother legal discussion somewhere down the road, but I think the data validity data bias data acquisition is a huge Achilles heel, we talk about, you were talking about LLMs and where it gets its data from.
And you know, the question is, it gets real hard if we have generative AI going, I don't know the answer so to make up the answer, do we have a way to find out if that answer is really true if we scrubbed it for me HRs that we can't really see. I think we have some huge transparency validity bias issues. I don't know, I can't wait to get to the other Achilles heel, which I think is a bigger one, but, you know, I don't know, Rayon, what do you think about that? I mean, I don't mean this is a tag we have a centralized data center yet.
We don't. And I mean, sort of bring it down to I think what people see and we're starting to notice more it almost is becoming a recursive problem. What I mean by that is every article you're probably reading now has been filtered through an LLJPT or cloud, right? So it's producing a lot of this content.
Now, what are these new alums being trained on? Well, they're being trained on the new content that's coming out that was produced by chat TPPT. So it's it's training itself on its own outputs and that's exactly what synthetic data is. And so you imagine just extending this thought forward maybe a couple of years.
The vast majority of data that it's ingesting is actually its own output. So what does that mean for the validity, the accuracy, the humanness of the model? I mean, you could sort of already tell and I think next week or coming up, we're going to be talking about AI slop. This is this is something that is becoming a real thing.
It's an, you know, the end of data for clinician. The end of real data. And this kind of sort of actually ties it back to actually the news of the day. This is where the I quote, the I know pun intended the idea of idea from top on is so powerful because it empowers clinicians and people are synch patients and researchers to, you know, work with the data that they're that they're generating.
And that data is super important. So like like we say is the new new oil, right? And so there's nothing synthetic about that. But yeah, I agree.
This is a major Achilles heel and everyone is sort of talking, you know, and the AI researchers are frontier bit are talking about this data bottleneck problem, the end of data, et cetera. And so no clear answer to your other than, you know, this is something we need to definitely attention to. I don't, yeah, I agree. I don't think there's an end answer yet.
But I think it's a big problem. A lot of smart people need to put their brains in on and figure out there has to be a better way. Or, you know, I love, we'll talk about AI slop maybe next week, you know, is that. Slop in, slop out, right?
Because eventually you just become so biased by because it keeps learning from itself. None of us get smarter when we learn from ourselves, because we always think we're right. All right, well, my next topic and this is maybe for the audience, we're stepping a little bit away from AI and I care. And we're talking a little bigger concept about the other Achilles heel of AI.
And it may not be what you think it's energy is that these generally these LLMs require so much energy. And we don't ever think about that. We're like, yeah, we just turn on this electricity. We're going to be good.
You know, I don't know where it's coming from. It's coming from wind farm or solar farmer hydroelectric or fusion fusion. You name it. But the reality is I did a little homework on this.
And there come out some interesting things. So training a single state large language, large language model like GPT three, not four, not four and a half, but three is equal to just training it is equal to almost 200 homes of electricity per year to train one model. And this is three. So when you get the four, which was what, 100,000 times more data.
I mean, I don't know if you can translate to 100,000 times, but now you're talking the energy of a city to power these LLMs. But even bigger than that is these data centers where it's all being housed because that's where the real energy is the amount of energy you run all the CPUs is absolutely crazy. And so data they just predicted the AI data centers have projected that the energy needed will double by 2030. And the data required the energy required to run all the data centers globally will be three times greater than the energy used in Japan in a year as a country.
That's a lot of energy, right? And you talk to some of the people I was telling Rayon in the pre-meeting out of guy coming to the data is AI infrastructure and he said our single biggest problem is in the United States is we're going to run out of energy by 2030. There won't be enough energy to power the AI systems and run electricity for our houses. And that was mind blowing to me to think that's a lot of energy.
And that's when you start thinking about that is, you know, for any of you who are kind of listening to the I news all the big tech companies like meta and like Google and Amazon are talking about, hey, we're going to put our data centers in space. We're going to move them out of land out of the firm at era and we're going to put them in space because space is cold. And these GPUs create a ton of energy, a ton of heat and it requires a massive amount of water to cool all of these systems. And then the energy we lose so much energy solar energy when it crosses into the atmosphere and they're like, no, we're going to put solar panels in absolute zero space and data centers are going to require tremendously less energy and then we'll just port it down to the US by or port it down to the rest of the world by starlink or whatever method that they're going to use to do this.
It was a very novel thing when I started thinking about the other day going, you know, the way they're going to move data out of the out of the world is they're going to move it out of the world and they're going to put it in space. That's a crazy thought is that we're not going to send astronauts to space. We're going to send astronauts to go fix the data centers that that we all depend upon. You know, and you know, maybe Elon Musk was not that far off with with starlink and think that's going to be the way not just we communicate with each other, but the way AI communicates with AI.
So I think about the energy issue and go, you know, when I hear people who do this all day say, there's just not of energy we made on the planet to support AI development in its current stage, much less having it grow to you know true artificial general and the way they're going to do it. I don't know how much energy that would take right I don't even have any concept of that but it's crazy when you think that our weakest LinkedIn AI. Is that you won't have to worry about pulling the plug because there'll be no more energy coming through the plug. No, that's that's that's pretty wild to think about so every time you're on using one of these elements think about all the jewels you're using in kilowatt hours to be asking you know how the weather is and etc.
So you know, I listen to Texas where we this is something we talk about a lot with my friends or an energy and there's not going to be a dispassion I've not heard of that's fascinating. There's going to be multiple prongs you know a lot of these tech companies like meta and Microsoft actually buying and repurposing old factories energy, energy factories and so nuclear is on the table. Of course solar and wind and other renewables I think are on the table and so yeah creative solutions are abounding this is a major problem that is definitely coming up so I definitely agree with you Scott the two Achilles heels are data and power and a lot of tensions being given to it so it's good for our audience to know about these two major issues that are I'm sure keeping a lot of people up at night. Yeah, I just think when you hear about the news and you hear about hey these companies are doing this we have to think why are they doing this right I mean why do they want to build a data center in space mean what's wrong with the data center here and then you start putting the pieces together and go.
There and go holy cow that is pretty much crazy and I just did it why we're doing this why we're talking why you're doing your rebuttal on that you know I looked up how much energy does one single AI powered search it is equivalent to you are heeding your microwave for one minute. Wow a simple GPT search or Gemini or whatever is equal to one minute of microwave that's the energy level. So I think about how many queries I do in a day I mean I do hundreds. The wonder my energy bills so I'm just kidding but you know I just it's a crazy thought you know we we kind of just my my son you know ran we type my son's Electroengineering is dad you don't need to know how to turn on turn the light switch and all the electricity that's my job just turn on the light switch and hope the light works.
Yeah, you know I think many times we do a search and we just think well okay this what it is but if you look at the backside carbon footprint and all the other stuff that goes with it it's a big deal, but I think a is expensive. AI is expensive and I think coming we'll talk about over investment in AI that's a topic we definitely need to discuss the bubble the bubble the bubble the yeah yeah the bubble. Everybody. For the audience hope you guys learned something I really learned a lot about the R.A.G.
is the rags I'm so happy you did that hopefully we could use some things to think about in fireside chat today about you know hey every time we do it it has a cost and what's the data you know can you trust the data so. We need to do is listen to us every week we love doing these we have some guests coming on in the near future here which will be really exciting. Thank you for joining us on real talk number 26 and we look forward to have any next week. You've been listening to real talk an AI and I care your weekly podcast to keep you informed about AI technologies revolutionizing I care.