Dylan Patel
Hello everyone, welcome back to SemiAnalysis Weekly episode number 19. We're here with the Steel team—one of the coolest things, actually, definitely the coolest thing that SemiAnalysis is doing right now: the Steel teardown lab. We're going to dig into the public launch of Steel. That means an article where we did a teardown of some consumer chips and put on display everything that the Steel team has to offer. Joining me today, we've got Andrew. How's it going, man? Welcome to the show.
Andrew
Thanks for having me.
Dylan Patel
And we've got Afzal. Welcome.
Afzal
Thank you. Hello.
Dylan Patel
Awesome, guys. I think a lot of the listeners are going to need a basic introduction, so hopefully you can bear with me as we go through this and explain some things to the general audience. Let's start with this: Andrew, can you tell me what's the Steel team, and what is a teardown at a high level?
Andrew
Yeah. What is a teardown? It's literally what it sounds like, right? We take, in the case of our first article, a consumer chip, take it out of the package, and start breaking it down and looking at what's there. We take it from the box we get it in all the way down to the transistor and everything in between.
Anyone listening understands that data center, AI, and consumer chips—everything—is advancing at the speed of light, in some cases literally. There's a lot of competition within the marketplace: who's in, who's out, what are the advances, and what are the technical nuances of every technology? Competitors are curious about how their competitors, within the same boundary conditions, are making decisions and advancing their technology.
A teardown goes into every aspect of that: materials, process, integration, electrical engineering, architecture, design—every piece of that puzzle that we can look at. That's what we're spinning up this Steel lab to do. Our first article, which came out a few weeks ago, is the first taste of what's possible and what we're capable of.
Dylan Patel
Awesome. It was a great article, and we'll definitely dig into it. Before Steel existed, when it comes to teardowns, people would obviously use the outputs that Steel can produce to understand chips. Can you explain a little bit about the motivation—what you use the output from the Steel lab to do in order to understand chips?
Afzal
There are really 2 angles you can approach this from. One is from the chip designers themselves, and one is as a competitor. For example, if I am a competitor and I have a teardown of an NVIDIA GPU, I can see how they're using the area, how much cache they're using, and how much area is being used for compute, memory, or I/O. That's one major thing for them.
On the other side, chip designers themselves get to see what the process node is before they even start designing. They can see a chip from Apple, and then they can see: this is what TSMC N3 looks like; this is what FinFlex looks like. Then we can say, this is our goal for when we actually make our own chips.
Dylan Patel
Awesome. Can you tell me a little more about exactly this article? Obviously, the title talks about SMIC N+3 as well as Huawei's Kirin 9030. So that's both the process technology and the chip itself that you might be analyzing. Can you walk me through a little bit of the high-level findings from tearing down the chip?
Afzal
First, let's go through some context. In 2021, SMIC started fabricating its own N+1 node, which was roughly equivalent to an 8-nanometer node. One of the main problems was that it didn't have any SRAM, so it couldn't be used for any big smartphone chips or AI accelerators.
Then, 2 years later, there was N+2, which they used in a lot of their smartphone chips, and they're going to be using it in their new Ascend G. Now we come to N+3, which is their newest one that they just started using for smartphones.
N+3 is still a shrink of the previous 2 generations, but the main thing is that, firstly, M0, which is the lowest metal layer, has shrunk a lot—by over 15%. Then you have the higher layers, which haven't shrunk as much. At the transistor level itself, there are so many major changes.
I guess one of the biggest things for their use case is that the SRAM is much smaller now. If it's even 10–20% smaller, that is much better for any new chips you make, because with some chips, maybe 2% is just SRAM. So it's a very major component.
Dylan Patel
Awesome, yeah. Okay, Andrew, how do you actually go about doing some of this stuff? If you're going to try to figure this out, can you walk us through the initial approach of sourcing the chip, and then where you go from there using all the incredible equipment and lab that you guys have built in Oregon?
Andrew
For sure. How do you get it? Consumer chips are easy. You go to your neighborhood electronics dealer and buy them; they're relatively accessible.
When you do a teardown, you need a number of samples. A lot of the stuff that goes on Twitter or wherever else looks easy: you get a nice, pretty picture and all these different details and analyses. But the amount of work that goes into that—the number of chips and samples that you need to actually extract all of it—can be quite ridiculous. That's what makes a lot of these consumer chips quite accessible.
When we start getting into the data center and other places, as you said, acquiring these chips and accessing them at those price points becomes a very different thing. It means we have to do a lot of planning. We don't just get these samples and go crazy with them; we actually have to plan this out.
You have to understand the technology and have expectations of what's there. This feeds into other parts of SemiAnalysis, whether it's VLSI or IEDM. We summarize all of these fantastic talks from manufacturers, foundries, and design companies. We need to understand all of that, understand what's going into a part, and that's just step 1.
So we get this thing in our lab. We have a plan, we think we know what's there, and we start unboxing it. In the case of a phone, we start pulling the screen off and looking at the chips inside. It's all pretty and fancy, but then we actually have to extract those samples.
Everything is soldered together, everything is packaged, and in the case of a smartphone, you have an SoC. How is that SoC designed? Looking at a domestic chip versus different competitors and all the OSATs that are out there, how are these things evolving? What's inside?
As we start looking at these, we can take a picture. That's fine, but we actually have to break them down. We tear them down, cut them up, polish them, and do all these things. It's very mechanical and very destructive.
Behind you in your screenshot is an X-ray. X-rays are a nondestructive technique, just like going to the doctor. You look inside. It's not the best for you or the part, but it's what we would call nondestructive, and it gives us an idea at the micron scale.
We always talk about nanometers and transistors, but at the micron and millimeter scale, there's a whole lot of detail and innovation there. Packaging is accelerating at a light-speed pace, much faster relatively than transistor technology is. That's not to say that there's not just as much work, and probably more money, going into it, but it's so far ahead in terms of complex packaging.
X-rays are an example of these tools where we can start looking at what's inside and getting an idea of what's there. But then we have to break it up. If we cut it in half, people love shiny die maps. There's so much you can understand from architecture, design, scaling, and layout. You can do those things, but once you have a die map, you can't cut it in half because you've already removed all the interesting stuff.
We have a lab where we can do these delayering techniques. We can reveal the die map, reveal the floor plan, cross-section, and cut. These are all very mechanical, hands-on things. You don't really think of that in high technology, but it's just like in the fab: you have CMP, chemical mechanical polishing; wet etch; dry etch; and all these advanced analytical tools.
We're looking at the most advanced technologies on the planet. Just like the people who develop those technologies need the most advanced tools, a competent teardown lab needs those exact same tools because we're looking at things at the same scale and complexity. This is really why there are so few players in this space. It's an intensive thing, and it shows the commitment of SemiAnalysis and Steel to make this a very strong and value-add venture—not only for SemiAnalysis, but for all of our clients and anyone who can read our free material as well.
Dylan Patel
Awesome, yeah. Okay, maybe we can walk through these things one at a time. The headline—the first picture out of the lab—is a die shot. Can you guys maybe just define what a die shot means and what it takes to get that first picture out of the lab?
Awesome. I'll have you explain really what it means, and then I'll take over.
Afzal
Yeah, sounds good. So, what we have is the die shot, and it's a full overview of the chip. On every die shot, you'll have certain structures and certain blocks. For example, you have a CPU core, a GPU core, an NPU, and your I/O. All of these comprise your floor plan, your layout of the chip. Then you can see that at a high level just with the die shot. You can see in the picture we have up there all of the big blocks and regions of the chip.
Dylan Patel
So, what does it take to get a die shot and then to get the annotation done to the quality that you guys are able to do for this?
Afzal
Andrew would answer this better.
Andrew
Yeah, this chip—this is the die, this is the silicon within the package. It's within the phone. We have to just break that apart. But in a mobile device, this die is actually embedded inside an SoC, a system-on-a-chip, right? DRAM, memory, interposer, BGA solder bumps, and then finally the silicon are all sandwiched inside of there. Through a variety of heating and mechanical processes—desoldering and infrared heating—we're able to extract the SoC from the phone.
We're able to start removing all those dies. We reveal just the piece of silicon, but the silicon itself isn't an entire stack of material. You have the transistor, metal 0, the front end of line, the interconnects, and the back end of line, all the way up to these relatively large structures. At that level, that's where you're actually contacting the chip. It's all the signal, power, and routing. But that really hides all the information that you showed. That floor plan is all the way down at the transistor level, where all of the SRAM, memory, logic, and other functional units are.
We actually have to work our way through all the metal layers. We have to get through all the back end, all the interconnect, down to what we call the poly. That's a bit of a misnomer these days. There's no more poly in these devices. It hails back to when planar transistors used polysilicon for their gates. These days, it's tungsten, tin nitride, and these other materials.
By actually removing material through chemical-mechanical polishing all the way down to that silicon, we can extract a whole lot of information, even at that nanoscale, using optical imaging. That's what you see here: these very high-resolution optical images from which we can extract a lot of information about different functional blocks and architectural decisions. That's really where we put a lot of labor into the lab to extract these. We hand that data off to the experts in architecture, design, and silicon layout, who can analyze all of the history, progress, and optimizations that all of these different design houses put into the fab and into a final product.
Dylan Patel
Quick sense check. You're kind of introducing another term, and obviously a die shot has annotation, but maybe you guys could explain a little bit about how the floor plan of this chip and the analysis can impact somebody's understanding of how a chip works.
Before actually getting a die shot and getting it annotated, maybe you have a certain understanding of how one of these SoCs works. After getting it done, you have a different, deeper understanding of it. Is there something that maybe you guys learned about this particular SoC, or is there a more generic point that you can make to explain why somebody would use a die shot like this to inform what more analysis you would want to do on a given chip?
Afzal
For this example here, the CPU and the GPU were relatively well understood because when you do your benchmarks, when you do your reviews, and even when the company announces it, they'll usually be quite clear on how the CPU works and how well the GPU works. But one thing we noticed was the NPU. In the previous generation, the 9020, it was only one slightly bigger core called a Lite Core and one Tiny Core. But in the new one, now it's one Lite Core and two Tiny Cores. This was something that we really didn't know before because, frankly, nobody knew how to benchmark an NPU, and Huawei hasn't said anything about its own NPU inside the Kirin SoC.
Dylan Patel
Yeah, so on screen, on the left we've got the 9020, and on the right is the 9030. You guys view this as a public contribution to the public's understanding of how this chip works, right? This is just us giving away some free information that we're able to understand based on the work that's done in the lab.
Afzal
Yeah, definitely. This is definitely something that wasn't known before, and you're just giving it to the public. It's not super private. I know that chip can't in theory get it out, so there are some people, especially the Chinese in China, that will buy the chips themselves and then do their own die maps.
Andrew
I think this is an opportunity, right? For us, for SemiAnalysis, a lot of these consumer devices are very interesting to people. They're a very different technology from what might go into a data center or AI, where we as a company also have a lot of interest. This is an opportunity for us to both demonstrate what we're capable of and reveal some quality information that's of value to our clients and our readers, and just put it out there. Give people a taste of what's behind the wall that we're looking at in the more data center and AI space.
That's an area, too, where we welcome any ideas or anything that's interesting. Bring it forward. We're up to the challenge.
Dylan Patel
Okay, so I'm really interested in what you learned about the SMIC process as well here. I'm going to skip packaging and memory comments on the chip for now, unless you guys have comments, and dig into the process. At a high level, to start this section, what's the takeaway, I guess, in terms of an understanding of the SMIC N+3 process that you guys got from this analysis?
Andrew
I think all of this is just taking a step back and looking at the big-picture things, right? It's a very interesting case study. Advanced leading-edge technology has moved on with EUV, and there are certain geopolitical reasons why SMIC is not able to use that. When one hand is tied behind your back, you figure out a way to innovate, right? You adapt.
This is absolutely an area where these restrictions have forced innovation, forced progress, and compromises that other fabs and other foundries may not have had to make. There's almost an analogy back to Intel 10-nanometer here, where they chose not to go with EUV. These same scaling challenges were present, and they had to innovate as well. Everyone saw the performance, yield, challenges, and delays that went on there.
But here we are with SMIC N+3, with extremely aggressive scaling using DUV at the M0 layer and above. I don't know if you have any of the pictures up, but we can look at that process. We can see where they're reaching parity with the world's leading fabs, and we can see where they're making those compromises. The metal 0, the different layers and etch stops, barrier layers, and the different metals are very clean.
But there are other aspects where we see that, with their final etch stop and their seed layer, and with the taper of the profiles of these metal layers, they clearly made some process decisions to manage yield and performance. With the right patterning, they might not have had to, but they found a way. They sell these in the market.
Dylan Patel
Okay. So, what's the next layer down in terms of the process? Of course, I'm thinking in terms of how you guys do the analysis when it comes to a teardown. You get a die shot, and then you start moving on to other things, particularly the TEM cross-section. Maybe you can talk a little bit about that and some of the tools that you use to actually do this analysis.
Andrew
Absolutely. We're really looking at the silicon at this point, right? Using TEM, the features are aggressively scaled down to the nanometer level. There's a variety of tools that we can use here. Again, we use mechanical polish, something that seems very rudimentary, but with the right technique and experience, you can actually reveal a lot of tiny structures and detail.
Going a step beyond that, we use a tool called a FIB, a focused ion beam tool, that actually uses gallium ions. You focus them into a tiny beam, and you can scan it across a sample and actually sputter or ablate the material away. When we have a floor plan, we know that there are different structures using different types of transistors or routing, so we can use this tool to dig in and create these cross-sections.
Those cross-sections can be imaged in an SEM, or scanning electron microscope. An SEM is perfectly capable of looking at things at the nanometer scale all the way up to the millimeter scale. But when you're getting down to the front end, when you want to look at a transistor itself, your gate metals, all of your contacts, and your interconnects, you need to go to a tool called a TEM, or transmission electron microscope.
And so, much like an SEM, you use a scanning or parallel beam of electrons. You accelerate these things to crazy-high energies, which makes them have very short wavelengths and allows you to resolve these tiny, tiny structures. But the crazy thing about TEM is that you actually can’t just image a face. You can’t just look at something like you would with your human eye.
You actually have to make these things incredibly thin—hundreds of atoms thick—in order to look through them with these electrons. The electrons pass through it, they interact with it, and on the other side, you can collect an image. That’s exactly what you see in these TEM images, which allow us to look at the fins, the interfacial oxides, gate metals, contacts, and interconnects.
Right? This is where we can extract the tiny functional units, the different types of cells, and the way the circuits go together. But also, how do you contact things? How do you route things? What different metals and materials do you use—conductors, semiconductors? It’s a whole world in the periodic table that goes into these things, a whole world of processes—very complex processes.
Thousands of steps go into making these devices: billions of transistors in a phone, across millions of phones. There’s just so much detail and nuance that we go under the microscope and look very locally to try to see how it’s done.
Dylan Patel
Okay, I want to pick one thing out. I found a few of these images really interesting when reading the article and trying to understand a lot of this stuff. If we look at something like cell height, which is on screen right now, and comments on the reduction from N+2 to N+3 in cell height, can you talk about what that means in a chip?
I think a lot of people may have a high-level understanding that 7 nanometers is more than 5 nanometers, which is more than 3 nanometers, but they don’t necessarily understand how this actually applies to something specific like cell height. Can you talk through a little bit of that and what the result is when you realize this at the end, when you have these images? Can you actually draw out how people are using this process technology?
Afzal
Yeah. The main thing—two main things—that you can still transition when shrinking is, first, your cell height. This determines how many fins you have and how many metal tracks you have between them. Then you also have the gate pitch, which actually contacts the silicon channel that contains all your transistors, where all your electrons go through.
When you have these two, you have a rectangle shape, a grid shape, that you can start laying out across the entire chip. Within that layout, you can have a maximum of 1 PMOS transistor and 1 NMOS transistor. If you just keep expanding, that is your basic transistor density. Then, on top of that, you have all the DTCO boosters, which maybe I’ll explain later. All of those DTCO boosters will add more to the density.
That’s at the highest level. But then, when you go down, when you shrink your metal lines, your resistance will go up because they’re smaller, and also because the liner needs to become relatively thicker. If you had a 20-by-20 structure and you shrink it to 10 nanometers, now your liner is twice as much, relatively. That will add a lot of resistance, and you’ll add capacitance. All of those reduce the performance of the chip.
All the modern fabs—all of the leading-edge fabs—have found techniques to improve that and make it less of a problem. That’s how, even on your 5-nanometer node from TSMC, AMD can still clock to more than 5 gigahertz. Intel, on their Intel 7 node, managed to clock more than 5 gigahertz, almost 6 gigahertz even. All those factors—the resistance, the capacitance, and all the parasitics—will affect the final chip.
Dylan Patel
Maybe you could talk a little bit about the concept of a library here. You’re obviously trying to reduce or improve density across different process nodes, and we can see that improvement as you go from N+2 to N+3 in these examples in the article. But there are different ways in which people can actually design a chip to use this stuff.
Maybe you can explain the concept of a library, and then what you guys realized in terms of what libraries are being used in the Kirin 9030 that was torn down.
Afzal
Essentially, a library is a group of all the basic cells that the designer wants to use, like an AND gate or even an adder, or some small block that’ll be integrated into a CPU, a GPU, or anything else. All these small blocks have dimensions like the cell height and the gate pitch.
For example, an inverter might take up 2 cell heights and 1 gate pitch. Those are the dimensions for that one cell. If you keep expanding, then in this library you have different options. One will have this cell height, and another might have 50% more cell height, but you can have more fins, so you can have more performance, basically.
On the other hand, you might want to go down in fins or gate pitch, and then you have a different library. This library is just a set of all the cells you can use as a designer. On N7, for example, there were 2 primary libraries: the high-density one and the high-performance one. This also carried over to N6.
The high-density one was used by most people, like Apple and AMD. On the other hand, the high-performance library was used on a few Qualcomm CPUs for the highest-performance CPU cores, like the big cores and the prime cores. Even after that, there was another specialized library for NVIDIA’s A100 GPU. They used an HPC library, where the high-density library was 2 fins, the high-performance library was 3 fins, and the HPC library was 4 fins.
There’s a huge amount of current they can pass through, so you can have much higher performance for your GPU, for example.
Dylan Patel
Can you comment on the impact that export restrictions have had, or really just the fact that Huawei has had to use SMIC? What sort of impact does that have on the chip designer and their use of libraries?
Afzal
The main thing here is that Huawei was one of the best chip designers before the ban. They were developing on TSMC N5 at about the same time as Apple. If you know Apple’s relationship with TSMC, that means a lot, really.
When they had to transition to the SMIC nodes instead, which are much worse, to be honest, they had to adjust and make use of every single transistor to the best of their ability. For them, at least, it all came down to architectural improvements. That’s how the 9020 is better than their older 9000. The 9030 is better than their older 9000, which was on TSMC N5. The chip designers really had to focus on transistor efficiency.
Dylan Patel
Is there a time for you to also talk about DTCO? It has a role here as well, right, on any given process, on any given node.
Afzal
Yeah, it does, but I didn’t really mention it because N+2 and N+3 are about the same for DTCO.
Dylan Patel
Yeah, good point. Anyway, what jumps to mind when you hear this sort of discussion about libraries and stuff? What jumps to mind?
Andrew
What jumps to mind to me is what’s next in terms of the technology. I’m very much a manufacturing guy. I love the complexity of scaling.
There are 2 aspects of that. First, transistors are scaling, they’re reaching certain limits, and we’re having to innovate in different ways. Backside power is already here. Gate-all-around has already been here. For a teardown lab, those represent new challenges. They require new types of processing. That’s something you see with 18A.
Even with floor plans, you see that the quality of the floor plans and die maps on these technologies is much harder to achieve. That actually means there are all sorts of collateral challenges, not just in the design and development of these technologies, but also in terms of failure analysis, fault isolation, and debug. All of these other areas also have to evolve and advance in terms of technology, capability, and innovation. That’s the area that I live in. That’s where I find things very interesting.
At the other extreme is packaging. This mobile SoC is not the most interesting package, but there are still some interesting innovations there. We’ll be coming out with a few more articles looking at more complex packages very shortly. It’s the same thing: things are scaling to submicron for hybrid bonding and these other aspects.
Again, it’s pushing into these areas that are incredibly hard to manufacture and incredibly hard to analyze. For a teardown lab, that’s a lot of fun. It’s a new challenge. How do you collect, analyze, and see this information? Then you hand it off to the designers and the experts in that space, and they’re finding totally new things themselves, right?
Dylan Patel
Can you explain a little more about the challenge? What makes it more challenging to tear down for you guys? What’s the roadmap for you guys?
Andrew
Backside power is an interesting one. You have your top metals—you’ve always had your signal and routing above—but things are getting too crowded. There are all these different signals and different things that create crosstalk and have capacitance, right? That motivates trying to use the other side of the device to route power or signal, separate those things, and let them relax.
The title of this article was kind of a joke, right? It wasn't meant to be serious. But it really has to do with the backside power, which allowed Intel's signal to relax in terms of scale.
A lot of the work that we do for sample preparation, say for delayering, might involve coming from the top, removing all those metal layers, and landing on the transistor. Of all the challenging things there are to do, it's relatively simple. But now, all of a sudden, you don't have all of the silicon underneath. You don't have FinFETs; you have gate-all-around. If you take those same approaches, everything just falls apart.
You have to do things differently. You have to develop new processes. All the teardown labs, anyone trying to take these approaches, are having to figure that out. You really saw why it took a little longer for everyone to come out with those capabilities.
Gate-all-around itself is a new challenge, right? For every channel, you have the—whatever you call it—the MBCFET, RibbonFET, or gate-all-around. You have all these channels of silicon on top, surrounded by gate. They're effectively floating, in the sense that you have these different material systems.
If you're etching or processing things, you deal with the chemistry and these different materials and layers to try to reveal or remove them. All of a sudden, everything is no longer anchored to hundreds of microns of silicon; it's just kind of sitting there in some metal. That creates new challenges in and of itself for sample preparation.
Packaging, again, is really cool. We have these different imaging modalities. Optical imaging is limited to around a micron or larger, but you only see what's on top, what's transparent, or what reflects. Scanning electron microscopy and TEM can go down to the nanometer scale, even the atomic scale, but it takes so much work to look at things, and you're very limited in how much you can see.
In between, you might have X-ray, but X-ray really struggles under a micron. Now that packaging, hybrid bonding, interconnects, and all these things are scaling into this almost no-man's-land of technology and capability, how do you see the things that are in between what these different analytical modalities can do?
How do you approach that? How do you make sure that, not only as an R&D lab, you actually make the stuff mature and excel and turn it into a high-volume process? Then, for a teardown lab like ourselves, how do we get in there and derive all of the information that creates competitive advantage for our client?
Dylan Patel
Exciting, man. Also, on the other side, let's say you guys have done some teardowns and you know what's coming. What's most exciting for you? Is it the consumer stuff, the data center AI stuff, CPUs, GPUs, or switches? What are you excited about that's coming?
Andrew
Generally, it's the data center—all of the data center GPUs with the huge packages. For example, cores is up to 5.5 vertical going even larger. It's coming out in some new GPUs, so all that is very interesting on the packaging side.
Afzal
But one more interesting thing is on the client side instead. Huawei recently announced its logic folding, where you have 2 chips, stack them together, and treat them as a single chip because of the very small-pitch hybrid bonds. The first generation has super 1.5 micron hybrid bonds. Currently, AMD, for example, is only using 6-micron hybrid bonds in its V-Cache and MI300 series.
When you shrink it so much, it's effectively able to act like a single big circuit. You can join blocks on 1 layer with another layer, and it's relatively efficient. Well, that's what they're claiming, at least. Hopefully, we'll get it later this year, and then we can tear it down and see all of the amazing innovations that Huawei is doing in that regard.
It's basically their approach to avoiding DUV scaling even further. N+3 is already quite difficult, so if you had to go even further, maybe you'd hurt yields and cost a lot.
Dylan Patel
Excellent, guys. Is there anything you think has been left unsaid so far?
Andrew
Number 1, this is just the beginning. We're excited to share more content on the front page. We've got a lot going; we're cooking in the background as well. If you think there's some unmet need, we're very up to the challenge. We want to answer those unmet needs and solve the problems that aren't being solved elsewhere. We want to answer the questions you might have.
We're excited. We're hiring. We're growing. Reach out.
Dylan Patel
Yeah, it's a big team already, but definitely growing. Also, how about you?
Afzal
I would just say, look out for all of our amazing stuff coming out soon. We have a lot of consumer chips coming out, and we'll be glad to share what we find on all of them—all of the most leading-edge stuff and some of the most interesting advanced packaging. It'll be very good for everyone to see it.
Dylan Patel
Awesome, guys. Well, congrats on the launch. I'm excited to see more, and thanks for stopping by and sharing a little bit this week. Nice show. Thanks, Sharan.
Thanks, everybody for listening, and yeah, take care. See you on the next one.