ThinkEnergy
Avsnitt

Summer Rewind: Grounding energy: How to scale cloud computing and data centres with Cerio

Dela

Summer rewind: When we say 'the cloud' what we mean is 'the data centre'. Globally, data centres are projected to consume over 1000 terawatt hours in 2026. What does that mean for energy production, distribution, and consumption? Guest Phil Harris, Cerio President and CEO, joins thinkenergy to shed light on something we all rely on but may not fully understand. From efficiency to sustainability, environmental concerns to Cerio's role improving how data centres manage energy. Listen in for the future of cloud computing.

Related links 

 

To subscribe using Apple Podcasts: https://podcasts.apple.com/us/podcast/thinkenergy/id1465129405

To subscribe using Spotify: https://open.spotify.com/show/7wFz7rdR8Gq3f2WOafjxpl

To subscribe on Libsyn: http://thinkenergy.libsyn.com/

---

Subscribe so you don't miss a video: https://www.youtube.com/user/hydroottawalimited 

Follow along on Instagram: https://www.instagram.com/hydroottawa 

Stay in the know on Facebook: https://www.facebook.com/HydroOttawa 

Keep up with the posts on X: https://twitter.com/thinkenergypod 

---

Transcript: 

[00:00:00] Trevor Freeman: Hi everyone, and welcome to the summer edition of Think Energy. As a reminder, we are in podcast vacation mode, and while our normal day-to-day work continues, we are on a brief pause from our regularly scheduled episodes to recharge, rethink, plan for the fall. But we don't want to leave you without anything to listen to. So, we've pulled some of our favorite insights and episodes from the past year related to the energy transition and where we are. So, welcome to episode two of our summer rewind, the August episode. So, as a reminder, this summer, we're looking at that collision between digital infrastructure and physical infrastructure, how the world of evolving technology is shaping and influencing and changing the way that we interact with energy and, in fact, interact with the energy grid as well. Last month, we looked at my conversation with Lynn Pettis from Overstory and how Overstory and Hydro Ottawa have partnered together to use satellite imagery and AI to protect our grid. In part two of our summer rewind today, or the second episode, we're looking at the world of data centres, and we're hearing lots about data centres as AI grows and takes kind of more of an ever-present role in our day-to-day lives. And I had a conversation with Phil Harris from Cerio. So, we're going to dive back into that episode and look at kind of the staggering energy footprint of AI data centres, of the cloud, etc., and how this boom is really changing how we do data centre design and really forcing a change in how we do data centre design just because of the sheer magnitude of energy required here. And we're going to look at kind of what this means for the future of global power distribution. So, sit back, relax, hopefully somewhere cool in the heat of the summer, and listen to this episode of Summer Rewind with Phil Harris from Cerio.

[00:02:18] Podcast Intro: Welcome to Think Energy, a podcast that dives into the fast-changing world of energy through conversations with industry leaders, innovators, and people on the front lines of the energy transition. Join me, Trevor Freeman, as I explore the traditional, unconventional, and up-and-coming facets of the energy industry. If you have any thoughts, feedback, or ideas for topics we should cover, please reach out to us at thinkenergy@hydroottawa.com.

[00:02:49] Trevor Freeman: Hi everyone, and welcome back. Data centres have come up a number of times on this show, and for very good reason. They have become a key underpinning technology for so much of our lives. Every time we pull out that phone from our pockets to pull up directions or buy something online, or doom scroll on your social media or news site of choice, every time you use your phone, stream a movie, leverage an AI model, whatever you end up using your phone for. It's funny, as I read this list, I'm sure there's like some university student out there who's thinking, man, what is this old man talking about? We don't use our phones for that. Whatever the kids are doing these days, whatever we're doing these days with our phones, with our computers, our tablets, etc., all of that leverages infrastructure that most of us have never seen, and quite frankly, probably don't really understand. We talk about the cloud like it's this amorphous, nebulous thing. But in reality, we're talking about real hardware in a real building that uses real energy, mainly electricity, a lot of water. And this isn't really new. Like, we've been leveraging centralized data centres for many years now. But what is changing is the scale of the data centres that we're seeing now and the pace of growth in computing power that we need to do the things that we want to do and that our data centres are able to deliver. So, just to throw a few numbers at it, the traditional data centre servers that maybe powered the early days of on-demand online streaming services, for example, they used anywhere from 5 to 15 kilowatts per rack. But modern server racks that are used to power AI searches, for example, can hit anywhere from 60 to 100 kilowatts per rack. This is great from a power output per rack perspective, but it means massive energy needs. And that is showing up in the size of load requests that we're seeing from new data centres. New data centres today are asking for service connections that are orders of magnitude higher than those built even just 5 years ago. Globally, data centres are projected to consume over 1,000 terawatt-hours in 2026. And just a quick kind of refresher from high school or wherever you would have learned this, a terawatt is 1,000 gigawatts, which is 1,000 megawatts. So, 1,000 terawatt-hours, which is roughly equivalent to the annual electricity demand from the country of Japan, an entire country. So, given all of this, there are a lot of incentives to find ways to maximize efficiency and reduce some of that energy demand. And that's where my next guest, Phil Harris, and his company, Cerio, come into play. I'll let Phil get into the details of exactly what Cerio does, but essentially, their goal is to reimagine the data centre to maximize sustainability and reduce energy needs. Phil is Cerio's President and CEO, and has been in the networking and data centre industry for over 35 years, including at well-known companies like Intel and Cisco, to name two. And I'm really excited about this conversation, one, to understand how do we make data centres a little bit more efficient, or maybe a lot more efficient, but also just to really understand, like, what are we talking about when we talk about a data centre? What is actually happening, what is physically inside these buildings? And we'll get into a little bit of that in our conversation. So, Phil, welcome to the show.

[00:07:05] Phil Harris: Well, thanks, Trevor. I appreciate it.

[00:07:07] Trevor Freeman: So, Phil, obviously, we're here today to talk about your work building sustainable data centres, or trying to make data centres a little bit more sustainable. But before we get into that, you know, you've spent your career, you know, decades of your career at different tech giants, let's call them, Intel, Cisco, to two to mention. You've seen quite a bit of change, no doubt, over your time. Has that change, like does this industry change linearly? Does it grow fairly steady, or is it kind of big jumps? And are we on the cusp of any major shifts? What can you kind of tell us about the future of this sector, data, tech, etc.?

[00:07:54] Phil Harris: It's interesting. I think as companies start, and I was at companies like Cisco, for example, when it was a very small company to where it was a very large company, and this should be no surprise to anybody, the bigger the company gets, the harder it is to change. And they really find that the only way they change is when they absolutely have to, not because they want to. And that's a combination of just inertia and shareholders' expectations and a whole bunch of things. So, I would say that the bigger the company is, the harder it is for them to react. And so, I think small, nimble companies tend to do much better when there's a lot of transformational technology and development and changes in the overall ecosystem we live in. I think your, the second part of your question, you know, I look at the current situation as a point in time where a lot of companies will have to make some significant changes simply because we are hitting two new walls: technological walls, commercial walls, geopolitical walls, that are really sort of confining what people can do. So, I think what's about to happen is we're about to see a significant change. And this is not atypical in the industry. If we think about back into the start of what we would think of today as computer science around mainframes that were happening in the '60s, you know, for about a decade and a half, two decades, there was a lot of dominance around a particular way of doing things. And then some new innovational technology came along that rapidly changed that, scaled out, and it went from a very dominant set of players to a much larger number of smaller players who could then provide more innovation and more scale and more choice. And I think we're about to see that transition occurring as well.

[00:09:47] Trevor Freeman: So, is this, is there sort of like an analogous time, 10 years ago, 20 years ago, are we on the cusp of like the big, the big change that we've seen before? Like, what would you compare this to, you know, in the last 20, 30 years?

[00:10:01] Phil Harris: Yeah, I mean, I think there's been eras of compute. And if we say, I mean, we can find analogies outside of the compute world, but let's just stay in the computer science world. I gave the mainframe example as one, and then we went to what we call client-server, which scaled out rapidly. Telephony, we went from large, big telephone exchanges that started in the government space, went to very large organizations. Now, basically, we've completely scaled out how we make phone calls, to use that now 20th-century terminology. Nobody really makes telephone calls anymore. And we went through this with cloud computing and the internet, where there was a change in the approach to the way we did things that suddenly gave us a scale-out mentality rather than a scale-up mentality. And I think that's what we have to key in on here, is that we can, someone, I was on a panel yesterday where we were talking about scale. And I said, well, to scale or not to scale, that is not the question. It's how do we scale? Do we continue to scale up, which is the current model, or do we start to think about scaling out, which is a more distributed model? So, we go from a small number of big things to a large number of smaller things. And typically in computer science, whatever you want to, storage, compute, memory, telephony, everything we've ever done goes through this arc.

[00:11:32] Trevor Freeman: Yeah, it's interesting, and there's, obviously, my brain's going to immediately try and find those similarities between my world that I live in on the energy side of things. And it's the same question. There is no path where we're not expanding the amount of energy we need, we're not going to be using more energy. But there are different ways to do that. And there are different paths we can take: the business-as-usual, the just grow, grow, grow, centralized energy production and large-scale transmission, or there's a combination of like grow those things, but also find alternative methods, more DERs, more sort of like close to consumer energy sources and storage, etc., etc. And people that listen to this podcast know I kind of go on ad nauseam about this. So, lots of similarities there. Another kind of framing or foundational thing that I want to talk through before we really get into the meat of our conversation is helping ground both myself and our listeners in what exactly we're talking about here. So, we all use, whether we know it or not, we use, you know, like cloud computing constantly, whether it's in our calls, how we're using the internet, using AI more frequently now. What is the physical reality behind that? What's actually happening? What is, you know, the term "data centre"? What is a data centre for our listeners here? What does that look like?

[00:13:17] Phil Harris: Yeah, let's start there. And that's a great question. We started recognizing that the amount of power and space required for computers in companies and government and all sorts of different applications was getting larger than we could put in a room in a closet near maybe where people were using it. We had to start to create dedicated space, because the power requirements, the cooling requirements, just the noise—you can't hear this, but just in my basement, I have a few different compute systems that my wife continually tells me is keeping the neighbors awake. The reality is the environmental aspect of these things became very difficult. So, we created these purpose-built locations that had then different requirements in terms of access and facilities and power and cooling and staffing. And so, they became a new way of thinking about building compute infrastructure at a building level, not just at the individual computers themselves. So, a data centre is usually a very large room, or building, I should say, that houses large amounts of compute and storage and other networking equipment. There's a whole range of different technologies that go into a data centre that allows us to process information. That's what a data centre is. To give you some analogies, in the US, there's about nearly 6,000 data centres, depending on how you measure a data centre. In Canada, we have about 400. In Europe, there's about 750 that we can identify as standalone data centres. You can probably find more places where computers are outside of people's homes, but that's about the ratio we're looking at.

[00:15:10] Trevor Freeman: And we're seeing, I think, and tell me if I'm wrong here, like all this talk about the AI proliferation, data centre proliferation, we're seeing an expansion of these. Is that we're seeing the size of these data centres expand, or we're seeing just more of them popping up? Like, what does it mean when we say we're seeing like data centre growth because of AI? What does that mean?

[00:15:36] Phil Harris: Well, it's fascinating because now our worlds collide. Because the way we now think about how to describe a data centre isn't in the square footage or the number of computers. It's in how much power it consumes. And we now measure it in megawatts. It starts in 10 megawatts, or single-digit megawatts for very small data centres, into average-sized data centres in the tens of megawatts, up to now the hundreds and the gigawatts of consumption that you look at these hyperscalers. But I think we have to put this into a sort of a human scale. It helps us to put this in human scale. If I were to go back to ChatGPT, actually about now 15 months ago, ChatGPT-4, if you were to put that data centre footprint into the province of Ontario, for example, where you and I both are right now, it would be the equivalent of a million internal combustion engine cars driving 30 kilometers a day. If you ever drive up the 401, you probably don't want to see another million cars on the 401. But that's the amount of energy that we can think of in terms of a data centre of that scale.

[00:17:02] Trevor Freeman: Yeah, and again, putting it in the electrical industry's terms, what we consider as a large load—so, we have a specific designation of a large load request—that is anything 5 megawatts and higher. Up until recently, we would get one or two of those every once in a while. Like, it's pretty rare to get a large load request. We are seeing large load requests coming in at a near constant pace now. Like, the number of large load requests we're getting, and a lot of it is because of this—not all because of data centres or anything like that, but a lot of them are certainly driven by that need for more computing power, more facilities that support that.

[00:17:53] Phil Harris: That's right. And at the same time, we're seeing a demand on energy around now home EV charging and other aspects of the general distribution of power. Everything is taking a step function. But if I could just say one thing to your point about 5 or 10 megawatts was a high load, I think we may need to change that scale. It's almost inefficient to build a data centre unless you're somewhere above the 10 megawatt range, because at that point, get somebody else to do it for you.

[00:18:23] Trevor Freeman: Ah, interesting. Yeah, and that's where sort of like, almost like renting space in a data centre for a request of that size. Interesting. Something that, you know, I've seen kind of in your writing on your blogs is the idea that traditional data centres are really built for peak capacity, which absolutely mirrors the power industry. We build our electrical grids for peak capacity. And obviously, that leads to a fair amount of inefficiency. So, if you're building just to peak capacity, if you're not at peak capacity, there is an inefficiency happening there. Something that you identified, it's a stat from your research, talks about graphics processing unit usage rates as low as 20% or 25%. So, I'm assuming that means kind of like three-quarters of that hardware is sitting idle or not being used valuably. Tell us a little bit about what Cerio, what you're doing, what your composable architecture specifically is doing to reclaim that wasted power and cooling capacity.

[00:19:42] Phil Harris: Yeah, and so it starts off with the premise you correctly raised, is that if we think about the equipment, the physical equipment, and how we put these devices and these components together in a data centre, the same model we've been using today is about 30, 35 years old in terms of individual compute systems where we run applications, software, that has memory and central processing units, those typical things you have in a laptop or you have in every computer. But then we put these accelerators, these GPUs, companies like Nvidia now are the one of the most valuable companies on the planet, if not the most valuable company on the planet, because that's the technology they develop. But we're trying to put these new class of accelerators into an existing compute model, which wasn't designed for this. So, that in itself now starts to fragment the ability to leverage those resources in a data centre. And as you accurately said, you know, it's interesting, if I could geek out on this a little bit for the energy consumer in the room...

[00:20:53] Trevor Freeman: Please do.

[00:20:54] Phil Harris: ...we think about the notion not only of the megawatts of power going into the data, but we think about what we call power usage efficiency. And that basically says, whatever the power delivered to a data centre, how much of that is applicable to the IT systems in that data centre? A good, well-run, efficient data centre is about 1.2. That means that about 1.2 times the amount of power that's used is delivered. Your home, for example, is about 30 times the amount of power we use is delivered. We are very inefficient from our home use, by the way. But that's another problem to solve another podcast. But in this case, that's all true until we then ask the question, but what's actually being used of that equipment? And that's now in that 25% to 30% range at any point in time. And we refer to that as stranded and idle assets that, for whatever reason, aren't where the application is or aren't applicable to be used for the application at that moment, because they're in some other box. Or it's a time of day when people use equipment. And by the way, equipment like that isn't being used 24/7, but it's drawing power 24/7. So, there's lots of inherent inefficiencies in that model. So, what we do is we provide the ability to dynamically have pools of resources where we can dynamically attach resources to a compute system as required, at the scale you required, and allowing you to be much more efficient in the timing of that and the amount of equipment required to meet your end solution. And by doing that, we can increase the number of accelerators that you apply to a compute system, which inherently means you are much more efficient in those compute systems. Because it's not just the computers, as I said before, there's storage, there's firewalls, there's load balancers, there's networking equipment, all of that can now be much more efficiently used, all of that is drawing power.

[00:23:01] Trevor Freeman: So, is the idea then that the equipment not being used or when you're at a lower demand time in terms of computing power, you've got physical equipment idling, sort of in more idle mode, drawing less resources that you can then ramp up? So, the peak amount of equipment's still there, you're just being more efficient with it when it's not being used, and you've developed a way to sort of dynamically pull that in. Is that what I'm hearing?

[00:23:28] Phil Harris: Exactly. I'll give you an example. A data centre here in Toronto wanted to have a block of 128 GPUs they could service their customers with. With the current systems they were using previously to deploying our infrastructure, they had to deploy actually 200 GPUs and a very large number of servers to house those GPUs. By deploying Cerio technology, they brought that down to 136 actual GPUs, and they reduced the number of compute platforms by a factor of 4. So, they reduced it by 75%.

[00:24:10] Trevor Freeman: Wow, that's fantastic.

[00:24:11] Phil Harris: With exactly the same outcomes to their customers, with no contention for resources, no oversubscription of resources, just more efficient use of those resources.

[00:24:23] Trevor Freeman: Gotcha. So, still able to meet that peak demand, but not firing up that equipment when it's not needed.

[00:24:30] Phil Harris: Well, not just not firing, not having to have as much stranded equipment, because we can use all the equipment all the time.

[00:24:38] Trevor Freeman: Gotcha, okay. So, in when I was kind of setting up that last question, I used the term "composable architecture," and I'll admit that I pulled that from your material. Help me understand what that means. So, I've also seen you use "composable infrastructure," sounds a bit abstract. What are we talking about here? What does that actually look like?

[00:25:02] Phil Harris: When a consumer or someone who's building a data centre buys their computer equipment, they usually will actually buy the computers, the GPUs, the storage, and other things at the same time. And they will get delivered together, and that box now becomes a unit of compute capacity. But the thing about that is, whether you're able to use that entire capacity, the length in which that's useful, there's a lot of innovation churn right now as new things are coming through very quickly. But that box is now statically built for the rest of its life pretty much. IBM did a study, to take a server out of a rack, these big 6-foot racks or bigger where these servers are housed with lots of wires going into them, power and data and all sorts of things, it's about $1,000 a minute to take one of those servers out of the rack and either change something that's broken, update something. So, they just don't get taken out of the rack, because the average time to take a server out of a rack is about an hour. The math on that's pretty simple. So, if I'm spending $60,000 to upgrade a $20,000 or $30,000 server, I'm just going to leave it there and buy another one. So, that creates more of these stranded assets. So, composability says, let's separate these things into, as I said, pools of resources: compute, accelerators, and other devices, and have a fabric between them that allows us to in real time assemble a compute system that I need—that's the composing part—as I need it, because I can now take the resources anywhere in my data centre if you've got the right fabric, which we've built, that allows you then to real time build that compute system with exactly the same capabilities, exactly the same performance, and without having to change any of your software or the way the servers work. Everything has to be off the shelf to make this work, and that's what we've built.

[00:27:03] Trevor Freeman: Gotcha. So, two of the terms—and you'll forgive me, this is sort of a new sector for me—two of the terms that are used as metrics to determine performance are power usage effectiveness, and you've kind of talked about GPU usage. Is the industry moving more towards that GPU usage metric? Is that just something that you guys are kind of leading the curve on, or where are we at on that?

[00:27:36] Phil Harris: Oh, no, this is very much the industry way of describing not just efficiency, but requirements. And we use very weird terms for this. Every industry has their weird terminology.

[00:27:46] Trevor Freeman: Absolutely, yeah.

[00:27:47] Phil Harris: And we're now moving to the, for example, in AI, the number of tokens per second. When you and I put a request or a question into ChatGPT or Copilot or Claude, whatever we use, those words get translated into tokens, actually numbers. Every compute system is just a big calculator at the end of the day. We do massive processing on numbers. How many of those tokens can I put into the system? How long does it take to process those tokens and give me a response? And the tokens per second per watt is now what we're asking. So, how many tokens a second and what power per token is it costing me to process information? And that's the interesting way of thinking about how AI, for example, that's where you started this conversation, will be measured, is the most amount of tokens per second per watt. Now, right now, we're focusing on tokens per second. We're not looking at the last denominator, which is watts. So, that's why these data centres are getting so ridiculously large. And, you know, we even heard it in the State of the Union address in the United States earlier in the week where, you know, there's now the administration pushing cloud vendors and AI vendors to say, "Hey, pretty soon, you're going to be on your own about delivering power because, quite frankly, the way you're going, it's going to become untenable to think about that from a national grid perspective." Now, I think that may be a little bit into the future, but I don't think it's a completely unreasonable sentiment at this point.

[00:29:32] Trevor Freeman: Yeah, and I mean, you're talking about, and we talked earlier about the just the scale of energy usage here is reaching a new height, a new level. And if we break it down to the individual racks, you know, these these racks of servers or processors that you've got in your data centre, we're now talking about anywhere from 50 kilowatts to 100 kilowatts of cooling need. And that's the big driver of energy usage, I think, is correct here, is the cooling need per rack multiplied by, of course, big numbers to get those, you know, 5, 10, 20, 30 megawatt data centres we're talking about. When we talk about cooling and we talk about hotspots within a data centre, how does your approach differ from kind of the standard way of doing it?

[00:30:32] Phil Harris: That's a great question. And I think we should explain why the cooling part, it's a bit like buying really good, expensive Wagyu steak every day and then having to spend a lot of money on a gym membership to then go and burn off those calories. So, we put all this power into power these compute systems, but then we have to keep them cool. The faster they run, the more powerful they run, the hotter they get, but we need to cool them. So, there's this relationship between the more power we draw, the more cooling we need. And cooling is becoming, as I said, that sort of tradeoff for performance. Now, there's lots of exotic ways of cooling computer systems. We can just blow air across them, we can have liquid like the radiator in your car, or we can literally drop these compute systems into baths of solvents. Ferdinand Porsche, I like to use of other industry analogies. Ferdinand Porsche, the guy who obviously designed the first Porsches and the VW Beetle, realized if I could distribute the heat of the engine block with a horizontal block, I could blow air across it, it was much more efficient than trying to put a radiator to actually cool down the engine block the way that other cars who have the engine in the front. And it's because of surface area. Now, if I've got to put all my GPUs and CPUs and memory close together either in the same box or the same rack, that concentration of heat needs to be addressed with cooling. One of the ways we can address this is not only to be very selective when I compose the GPU, it's the only time it's drawing power, but also, I can spread them out through my data centre by having a fabric that allows me to connect them to the compute systems with the same performance. But now I can distribute my heat generation, that means I can cool more efficiently, just like that Ferdinand Porsche analogy of the Porsche 911. Because now, heat over spread of distance and surface area is a more efficient way, which means that it won't mean that we won't ever get to liquid cooling. I don't think immersion cooling is a good idea for lots of other reasons, but it's a necessity more than an optimization. But we can defer the complexity, the cost of those exotic cooling systems if we're more efficient in the way we use and design our data centres.

[00:33:10] Trevor Freeman: And I guess there's a similar description there of if you're concentrating all that heat in a specific physical area within a bigger building, room, whatever you want to call it, that cooling system is having to work to that peak cooling need, so to that hotspot effectively. But it's not working just on that spot, it's working across the whole physical area. If you're spreading that cooling need out across the whole room, one, the peak is a little bit lower, and you're just more effectively using your whole cooling system, is that fair to say?

[00:33:48] Phil Harris: That's exactly the right way of looking at this. And think about it from this perspective as well. The reason we have to cool is because if we don't cool sufficiently, those devices become very unreliable and reduce their useful lifespan. Without going into who, because they keep this information confidential, but one large cloud provider in the US, for example, a GPU that normally has a lifespan of at least 3 years is going down to about 9 months right now. And the reason for that reduction in the lifespan of the use of that GPU is because of the heating characteristics within these boxes even with all these cooling mechanisms are becoming now a reduction in the lifespan. So, that means we have to create even, remember I said what it costs to take a system out of a rack?

[00:34:44] Trevor Freeman: Yeah.

[00:34:45] Phil Harris: That means if we don't have to apply an efficient and effective cooling strategy, our power strategy and cooling strategy, then we start hitting problems very quickly.

[00:34:56] Trevor Freeman: Gotcha. Okay. Okay, so there's a mantra that I'll admit I hadn't seen before until kind of reading some of your material. It's, "Friends don't let friends build data centres." And I think it's referring to, you know, this move in there's so many industries that kind of do this cycle of centralization to decentralization. And the sort of data movement went towards that centralization, and you saw these big, massive data centres. But there's kind of a move now back to, let's call it, decentralization or repatriation of data. And so, for various geopolitical reasons, organizations, companies, governments are wanting to pull their data back home and have it kind of be more in their control, living in their own servers. So, how are you or how is Cerio helping companies kind of get back into the data centre business or repatriate their data without kind of, you know, getting into the troubles that led for to that centralization in the first place?

[00:36:12] Phil Harris: Yeah, and by the way, I can't take real credit for that quote. Cole Crawford, who was one of the early guys at Facebook before it became Meta, and was one of the leading voices in the Open Compute Platform movement, which was trying to standardize how we do these things, Cole is now the CEO of a company called Vapor IO. And what he was really saying is, it's so complicated and difficult to run data centres, let alone build them, the capital expense. AI isn't just one thing. There's lots of stages in the workflow of AI. We train these big models, you have heard of large language models like ChatGPT or Copilot. But what we use them for, the results of those trained models, is what we call inference. Now, you'll now hear about agentic AI, where we turn those results into actions. Okay, that's the agency part of agentic. Well, the use of AI in the corporate world is now becoming, as you said, both regulated, but from an intellectual property perspective, it's about how I control my data and my information. Because if I put that all into somebody else's large language model, I've basically populated somebody else's large language model with what might be my proprietary information or information that's very sensitive. And it's one of the reasons why you'll hear in the press about Anthropic, for example, trying to put guardrails around the use of their AI, because they're very sensitive to this. Most enterprises, governments, of all sorts, have realized, though, they need to run this in their own data centres, because they need to have control over this information and the use of this information. That's the repatriation you're talking about, moving these workloads now into the organization that previously had said, "Hey, cloud computing can take this problem." We've got to now figure out how enterprises, which are far many more of them in far more diverse locations, can now build their own data centres and get the right power, the right efficiency, the right capabilities at the right cost.

[00:38:22] Trevor Freeman: Does that open the door? I mean, earlier you talked about, you know, if we're talking about a 5 megawatt data centre, it's almost not worth it. You know, that's just sort of renting space in someone else's. How does how does that track with an organization that won't have enough data or enough computing power, whatever the metric is, to to warrant a 30 megawatt data centre for their own data, but wants to get that that control, wants to bring it more in house? Is your technology helping those smaller data centres exist? Is that the correlation there?

[00:38:58] Phil Harris: We can now move it into, another couple of terms that may be you're listeners may not be familiar with, in the compute world or the data centre world, we talk of brownfield and greenfield. Brownfield is that which is already there, greenfield is something I have to build new. A lot of the brownfield world is the predominant quantity of compute power on the planet is primarily brownfield. The question is, can I take that existing infrastructure and put the capabilities we've been describing in this discussion into those brownfields, so I can reduce the cost of the expansion of that, because I can reuse the compute equipment that's there? I can now add just the discrete GPU technology, for example, into an existing data centre that doesn't then far blow the power budget or the cooling envelope within that environment, but I can still now start taking advantage as I figure out what my larger plans are. And at the same time, how do we have a tier of providers who I'll give you an example. There's a company in again in Canada, ThinkOn, who are building a data centre in in Ottawa. It's going to have its own liquid natural LNG as its source of power for its own power requirements. Why? Because they can have the power they need as they need it in that location, and they can provide that secure infrastructure for both government and private enterprises. And ThinkOn is is certainly in Canada one of those companies that's really seen to be a trusted partner in this. So, it'll be a bit of what can I do myself, how do I have a trusted partner. We think of sovereign AI a lot. That means trust more than anything. And that's becoming the new mechanism of thinking about this.

[00:41:01] Trevor Freeman: Thinking about the the environmental impact of of tech and of of data, you know, we've talked about the energy usage here, but there's also the physical aspect to it of, you know, the the pace of improvement in technology means we see obsolescence, or we see kind of technology being outdated fairly quickly. We all, like on the personal level, we all see this with our our cell phones, our smartphones, our our whatever tech we have at home that seems to be out of date fairly soon. I think the the stat or the the saying that's out there is, you know, tech is kind of obsolete or becomes trash within 3 years. Obviously, this is not sustainable. Is this part of the drive of what you're doing? Is it are you looking to sort of extend the life of the physical equipment? You've touched on this a little bit, but maybe expand a little bit on that.

[00:42:01] Phil Harris: Yeah, this this goes a little bit back to that brownfield, greenfield discussion. But one way of looking at it, I guess, is um when I put all of these components into one the classic model, the current model, I put my my my central processing unit, my memory, my storage, my GPUs, all in the same box. What is the thing in that box that I want to take advantage of as new innovation happens versus that which is happening over a slower evolutionary cycle? Well, right now, if I put everything in the same compute unit, go back to my cost of taking that box out of the rack, I'm pretty much limited by the slowest innovation curve within that platform now as what I can take advantage of over time. Interestingly, GPUs are innovating currently at a clip of about once a year. Nvidia comes out with a new generation of GPUs once a year. Um but now we're getting more GPUs into the market, we're getting much more diversity, and that diversity means I'll have more options more often. But if my compute system itself is only innovating once every 3 years to your point, then if I don't decouple these things, if I don't have the ability to separate these innovation curves, I'm always stuck with the slowest innovation curve. One of the things we've done at Cerio with the fabric that we've built and the platform we've built is to allow you now to, if you like, dislocate those innovation curves and those options so as new technology comes along, I can apply it to the things that are innovating slower and still get the outcomes I'm looking for. And that will significantly increase the existing lifespan of equipment that's in people's data centres.

[00:43:56] Trevor Freeman: So, so looking at a data centre of the future, and not, you know, not far into the future, let's say 5, 10 years from now, are we seeing some of the same technology still exist within that data centre, or is it, you know, everything gets cycled out within Like what's the generation of a data centre, for example? Like how often or how soon will we see it all cycle out?

[00:44:20] Phil Harris: I think you there's a there's a there's a technical answer to that, and a financial answer to that. The depreciation models, so that the capital infrastructure can be written off people's books over a 3 to 5-year window, is very typical. So, we see that there's just an a financial inhibition to changing more or faster than that 3 to 5-year window. The technical churn, as I said, is happening much more rapidly in the technologies that are drawing most power, but providing most capability. So, one of the things we're we're looking at is how companies now start leasing infrastructure. Because if they lease the infrastructure, they can now recycle that and bring new technology in faster into their organizations. But to do that, you've got to have the ability to bring new technology in and not be stuck with these static systems that we have today. So, there's a set of financial instruments and now, with work that Cerio is doing, technical capabilities that allow customers to really continue to innovate. So, there's no real, "Hey, it's going to be all churned out in 3 years." I'll continue to innovate over those 3 years, recycling the technology that can stay where it is and bringing new technologies as it becomes available at the right financial model.

[00:45:39] Trevor Freeman: I'm curious about what that innovation is. Um so, you talked about Nvidia kind of essentially a new GPU every year, there's a new version every year. What is the innovation? Are they just Is it getting faster and more compute power and therefore it's pulling more energy, and is that just like a perpetual increase, or is it kind of same compute power, less energy? Like, do we ever see, I guess what I'm what I'm getting at with this little bit of a ramble here is, do we ever see that that rate of change in energy usage start to flatten out and come down while we still can grow our computing power, or does energy usage just continue to grow, like are we on a bit of a path with no end right now?

[00:46:36] Phil Harris: History taught us a little bit about this. Uh Gordon Moore, who was one of the founders of Intel, actually, we had this term called Moore's Law. And Moore's Law was basically this idea that every 18 months, we'll double the number of transistors on a piece of silicon. Now, for those in the computer science world, we all understand what that means. For the rest of the world, the transistor is the smallest unit of technology within the computer. It's the basic building block of how we build computers, the central processing units, all the GPUs, they all come down to taking literally silicon, and in a foundry, we call them, figuring out how to make as many transistors interconnect with each other in a in a smaller area as possible, or the most amount of transistors we can. So, a bit of a geeky answer to your question, but the way that we look at how each innovation improves is are we increasing the number of transistors, which means we can do more math. Remember, all we're doing is processing numbers.

[00:47:45] Trevor Freeman: Per unit, per physical unit, right?

[00:47:47] Phil Harris: Per physical unit.

[00:47:48] Trevor Freeman: Okay.

[00:47:49] Phil Harris: And the way we do that is in these big foundries that process all this silicon into these components, they have what are called process nodes. And the and literally how we etch a transistor, it's called lithography, onto a piece of silicon tells us the power of that piece of silicon. And the more I can etch, so we get into what we call the nanometer scale of what we call a process node. So, every time, if you really look into the spec sheets of Nvidia every generation, they'll talk about how many nanometers their silicon process is based on. Because the smaller I can get that number, the more transistors I can have on the same amount of silicon, the more processing I have, BUT every transistor takes power. So, with more transistors, I require more power, even though in the same physical space, it looks like the same amount of silicon. Therefore, your question was a great one. Do we ever get to zero nanometers? Well, no. We're going to hit a wall here eventually. So then the question is, that's the scale-up model: try and make one thing as big as possible. How about if we make lots of things powerful, but we have more of them? In China, last year, we heard of DeepSeek. DeepSeek was a Chinese government-sponsored effort to try and come up with a much more cost-effective way of doing the equivalent to ChatGPT. They didn't do that with bigger GPUs, they did it with much smaller GPUs, but many more of them. And that comes back to how efficient I am in deploying lots of things together. And that goes back to my earlier point about, we start with scale up, inevitably in the industry, we go to scale out.

[00:49:50] Trevor Freeman: And there's Is it fair to say that the power usage per transistor, is that fairly static? Like, is there efficiencies to gain there, or your GPU is going to use more power because you're packing more transistors into it, and once you hit that wall, that's going to be the the power consumption level. Is that right?

[00:50:15] Phil Harris: Well, this is the games that the silicon manufacturers like Intel, AMD, Nvidia, they're all trying to figure out how to sort of figure out new and interesting ways of packaging all the silicon in these processing units. And we've got a whole industry and science around the packaging mechanism to make those tiles, and we now think of them as little tiles of processing power. And some will be doing very specific jobs, some will be doing very general jobs. It's now getting to the point where the science around the packaging of these dies or these tiles is as much of the of the of the innovation as the actual tiles and the processing on them. So, it's an extremely complex um um technical problem, uh and we are hitting some walls here, which is why I go back to my earlier point. We're now reaching a point where is it just a technical problem to solve or a technical, operational, and commercial problem we have to think about? And this is that wall that you asked me about right at the beginning of this conversation: are we about to hit a wall? And the answer is yes.

[00:51:26] Trevor Freeman: Hmm, interesting. It's I mean, I'm always fascinated by like what are the what are the really smart people in the industry focusing their time on. And it's so that's why we're talking to you, um of you know, you're looking at how do we operationalize this, how do we get the most efficient combination and structure of what we're doing here. There's folks that are looking at how do we pack the most computing power efficiency into these specific units. I guess there's an aspect of how do we cool this in the in the most effective way, like what's how do we um, you know, drive down the cooling power needed. Uh what else is out there in terms of like we have smart people focused on this efficiency? What's the thing that's missing from that that sort of list?

[00:52:26] Phil Harris: Well, I think maybe what's going on right now, and if I could just add one more layer of complexity, I'll try and keep it concise. Remember I said we were processing silicon? Well, the earth's got lots of silicon, but we don't have lots of places to process that silicon. The companies that are formed to process silicon into these processing units, we call them foundries. The world's largest is TSMC based in Taiwan, um and then we have Intel, we have Samsung, we have a few others around the world, GlobalFoundries is another one. There is a limit, physical limit, because these foundries are huge and they take decades of development and optimization. So, if we start breaking ground on a new foundry tomorrow, we'll see output in about 5 years. So, we have a constrained supply. So, if I'm a if I'm Jensen at Nvidia or any of the big silicon manufacturers, I'm going to optimize that relatively constrained supply to where I'm going to get the best return on my investment, and that's why this scale-up model is happening. So, given that, we know that we won't have any more foundry capacity of scale for another couple of years at least, then the reality is we've got to think differently about how we're thinking about the processing of that silicon. Do I want just ever bigger processors that become more expensive, more limited in where I can deploy them, and quite frankly, the top 15 consumers in the world of silicon consume about 80% of that silicon, if not more. How do I democratize that? Again, it goes from scale up to a scale-out model where I can use that same processing capacity to produce more silicon.

[00:54:20] Trevor Freeman: Fascinating. Um yeah, I just I took us down a little bit of a nerd out path. You had me really interested in that. Um okay, so last question here. We hear this term for a bunch of different reasons um around the world right now. Um we're hearing this term "democratizing" happening a lot, and and I know um and you've talked about democratizing AI. What does that mean? What does that mean to you, or or describe that for us?

[00:54:51] Phil Harris: Yeah, I think it really means going right to my last point about if if 15 big, big consumers of silicon are going to consume the vast majority of available supply chain, that makes that a losing proposition for the the rest of the organizations and the rest of the governments and the rest of the individuals on the planet. So, how do we make sure that AI can be built both responsibly from a a sustainability perspective, and I don't mean just ecological side, but that's important here too, but also from the ability to I was on a panel yesterday between the UK government and the Canadian government where we were looking at how do countries around the world have the ability to control their own destiny. There's this whole notion of sovereignty and AI sovereignty right now. That isn't because people want to have closed walls around them, they want to have choice. They don't want to be dictated to by very dominant players where they quite frankly don't have the buying power to compete, you know, the amount of capital going into some of the AI companies, we saw $30 billion going into Anthropic last week, that's actually a small increase in their capitalization relative to the other big AI players on the planet. That's $30 billion. So, we've got to think to ourselves, is that a sustainable model commercially? And the answer is no. So, we've got to have technology, we've got to have the right ability to deliver power, we've got to have the right designs of data centres that can keep them cool in an effective and efficient and responsible way, and we've got to be able to give them enough power to make them viable to make them useful. That's the democratization we all have to be focused on.

[00:56:46] Trevor Freeman: And we need every I guess to to round out the point is we need everybody to be able, everybody being, you know, whatever, major industry, countries, whoever, to be able to access that equally so that we don't have to rely on the major players out there in order to do those things you just said. Gotcha.

[00:57:04] Phil Harris: That's exactly right. And look, there'll always be a pyramid here. There always has been in technology, there's always still the big players, right? But the question is, have the big players stifled out the ability for smaller players to come up, innovate, provide choice, provide alternative ways of looking at things? And that's what we've got to make sure that we keep the the And there's always reliance on some new technology coming along that enables that. Cerio believes that we've created that next layer in the stack, if you like, of technologies that gives us that opportunity to rethink the innovation curve going forward.

[00:57:42] Trevor Freeman: Very fascinating. Phil, thanks for your time. I really appreciate it. This has been super interesting. It's not an area that I often get to spend my time thinking about, so it was great to chat today. As you know, we always kind of round out our interviews with the same series of questions to our guests. So, what's a book that you've read that you think everybody should read?

[00:58:05] Phil Harris: Well, I'm not sure I can recommend this for everybody. One of the people who basically along the lines of some of the things I've been talking about today, who've revolutionized the computer world was a gentleman by the name of Linus Torvalds in Helsinki in Finland at the time, he's now based in the States. He realized that there was a dominance around how the operating systems on computers, the things that run the software, was limiting basically innovation, choice, and forcing us down a very closed path. So, he wrote something called Linux, which was a new operating system, so be it on your phone, your TV, your microwave, that's running Linux today. Because there wasn't an operating system that we could then generally deploy that meant there was more developers had the ability to write applications, more hardware vendors could now have software they could run on their on their platforms. He gave the world a new innovation curve, and every time this happens, to my last point, good things happen, very good things happen for the world, for every individual on the planet. And Linus was one of those individuals who saw that need. And so, his book, "Just for Fun," and he's a very quirky guy, as you can probably imagine, is a great book about his philosophical approach to what it takes to change really big problems. And I would encourage all of you just even just read the first few chapters. It's a fascinating view of how an incredibly smart man, smart individual took on probably one of the biggest problems we had in the 20th and 21st century of computing and solved it by recognizing you take a different path.

[01:00:03] Trevor Freeman: Yeah, very cool.

[01:00:04] Phil Harris: As far as as far as shows, I don't know, I'm one of these guys, I've got two 13-year-old daughters, so my wife and I get to watch TV for a very limited amount of time when we can watch it about the things we want to watch. So, we tend to sort of cram things in. But I am I am a huge Aaron Sorkin fan. So, if I ever need something on a rainy day to go back just to think about how the world could be, I watch The West Wing. It's a show that's imaginative, it's got incredible script writing, it's got incredible character development, but it really talks about how to think about doing the right thing as well. Now, whether you agree with the politics or not, that's a different question, but just the thought that smart thinking solves big problems, again, sort of it's a bit like the Linus Torvalds book, it just speaks to me about sometimes we can solve big problems with individuals or people who just have the right way of thinking about things.

[01:01:06] Trevor Freeman: Yeah, I think that's that's the kind of, you know, call it entertainment because it is entertainment, but it's the entertainment that sticks with you and that we we go back to time and again is the ones that we can also like see the the underlying philosophy or or, you know, theory of change that goes into that entertainment. And it's it's fun to watch, it's, you know, either humorous or dramatic or whatever, but there's still that underlying message. And I think, yeah, West Wing is a great example of of that. There's a handful of those other sort of classic shows that are in that line, too. Um, a free round-trip flight anywhere in the world, where would you go?

[01:01:46] Phil Harris: This is hard. Um, my wife and I were talking about this the other day, and um I've had I've had the luxury of traveling just about everywhere. I think there's 15 countries on the planet I haven't been to. Wow. But if I ever want to go to one place, it's Bali. And there's two reasons: one, my wife and I went there for our honeymoon and it was the beginning of the most important chapter of my life by far. Um, and and secondly, it's because it has that balance of everything. It's I love to scuba dive, I love the rainforests, the jungle, the architecture, the people, the food. It just brings everything into one package for me. And so, um it just, again, it's those things that sort of speak to you emotionally and also intellectually. It's one of those things that I could always go back to.

[01:02:40] Trevor Freeman: Oh, fantastic. Um, who is someone that you admire?

[01:02:45] Phil Harris: In history or today?

[01:02:47] Trevor Freeman: Ah, you pick, anything.

[01:02:48] Phil Harris: That's fascinating. Um, I think historically, it's I'm a Brit, it's hard not to go back to some of my my my forebears or my my country's forebears. Alan Turing, who against all adversity, social, political, technical, came up with an inspirational way of thinking about solving what were deemed to be unsolvable. And again, it was a it's a tragic story, I think we've all if you see, you know, the movie that was made about his life, um, it's a very tragic story, but it's an inspirational story about how again, if you just take a different approach to solving what seems to be an unsolvable problem, you can. You get smart people together, doesn't have to be a big army of people. And I think so Turing is one of those people that always comes back for me, thinking, "Wow, if I could have just some of his courage and some of his imagination and some of his intellect, I'd be a very happy person."

[01:04:02] Trevor Freeman: Yeah, and it's almost, I mean, obviously a brilliant man, but it's the willing to think in a different way or willing to approach a problem in a different way that, I mean, there's a long list in history of of major turning points that are as a result of someone thinking in a different way or doing something in a different way, and I think that's a great example of it, so. Just about the entire course of human life in the midpoint of the 20th century changed on that that man's inspiration, that man's imagination. Yeah, and that's that's not an understatement. That's fantastic. Uh, okay, last question. What's something about kind of the energy sector or, you know, your sector that that you're really excited about, or something that you see in the future that you're really excited about?

[01:04:47] Phil Harris: Actually, I see it now, to be honest. There are things in the future. Hey, I I I have two 13-year-old kids. I want to have a sustainable ecology and world environment for them to live in and bring their own families up in. And I think about how we can use power more efficiently, but how we can make it It does Look, sustainability is important. I want to see renewable, sustainable energy for the general world as a as a thesis. Right now, it's how we can be much more efficient in the use of power and the right power delivery. And I think, as I said, I gave the ThinkOn example, that's incredibly exciting because now if we can do that at scale, that's an opportunity to do that democratization that I spoke about. So, when I think about the things that are really exciting me about the data centre world, the world I live in, actually that power generation and power availability in a clean, effective, well-managed fashion is exactly what we need right now while the rest of us are solving these transistor problems.

[01:05:58] Trevor Freeman: Yeah, it's I mean, our listeners are probably going to roll their eyes because I say this all the time, but one of the things that excites me the most is seeing like we're in a period of change, and and that's a really exciting time to be working in this, and I kind of hear that from you in in your sector as well, and I see it in mine, in the energy sector, of we're we're actually getting to see some of this innovation, some of these like leaps and bounds forward. That's not to say there aren't still problems, that's not to say there aren't steps backwards as well, but it it's very cool to be working on this at a time when we're seeing that change, and and that's kind of what I'm hearing from you as well.

[01:06:34] Phil Harris: Indeed.

[01:06:35] Trevor Freeman: Awesome. Phil, thanks so much for your time. I really appreciate it. This has been great chatting with you.

[01:06:39] Phil Harris: Trevor, the pleasure was all mine. Thank you.

[01:06:41] Trevor Freeman: Fantastic. Take care.

[01:06:42] Phil Harris: Take care.

[01:06:43] Podcast Outro: Thanks for tuning in to another episode of the Think Energy podcast. Don't forget to subscribe wherever you listen to podcasts, and it would be great if you could leave us a review. It really helps to spread the word. As always, we would love to hear from you, whether it's feedback, comments, or an idea for a show or a guest. You can always reach us at thinkenergy@hydroottawa.com.

Podden och tillhörande omslagsbild på den här sidan tillhör Hydro Ottawa. Innehållet i podden är skapat av Hydro Ottawa och inte av, eller tillsammans med, Poddtoppen.