Lecture 1: Introduction
MIT 6.824 Lecture 1: why distributed systems get built, how the labs work, fault tolerance and consistency, and a walk through Google's MapReduce.
Transcript source: automatic speech recognition on Vidleaf (unedited, may contain errors). Paragraph breaks and timestamps added by Vidleaf.
[0:01] All right, let's get started. This is 6824, distributed systems. Um, Thank you. So I'd like to start with just a brief explanation of what I think a distributed system is. Thank you. The core of it is a set of cooperating computers that are communicating with each other over network to get some coherent task done. The kinds of examples that we'll be focusing on in this class are things like storage for big websites, or big data computations such as MapReduce, Um, And also somewhat more exotic things like peer to peer file sharing.
[0:43] Those are all just examples of the kinds of case studies we'll look at. Um, And the reason why all this is important is that a lot of critical infrastructure out there is built out of distributed systems. Infrastructure that requires more than one computer to get its job done or sort of inherently needs to be spread out physically. Um. So the reasons why people build this stuff, First of all, before I even-- talk about distributed systems, I just want to remind you that If you're designing a system, you need to solve some problem, if you can possibly solve it, on a single computer.
[1:17] you know, without building a distributed system, you should do it that way. And there's many, many jobs you can get done on a single computer, and it's always easier. So distributed systems, you should try everything else before you try building distributed systems 'cause they're not simpler. So the reason why people are driven to use lots of cooperating computers are They need to get high performance. And the way to think about that is they want to get, achieve some sort of parallelism. Lots of CPUs, lots of memories, lots of disk arms moving in parallel.
[1:54] Another reason why people build this stuff is to be able to tolerate faults. Wow. . have two computers do the exact same thing. If one of them fails, you can cut over to the other one. Another is that some problems are just naturally spread out in space. Um, You want to do interbank transfers of money or something. Bank A has this computer in New York City, and Bank B has this computer in London.
[2:25] you know, you just have to have some way for them to talk to each other and cooperate in order to carry that out. So there's some natural sort of physical reasons systems that are inherently physically distributed, And the final reason that people build this stuff is in order to achieve some sort of security goal. So often by, if there's some code you don't trust, or you need to interact with somebody, They may be malicious or maybe their code has bugs in it. So you don't want to have to trust it. You may want to split up the computation so, you Your stuff runs over there on that computer, my stuff runs here on this computer, and they only talk to each other.
[3:03] to some sort of narrow, narrowly defined network protocol. So we may be worried about security, And that's achieved by splitting things up into multiple computers so that they can be isolated. Um. The-- Most of this course is going to be about performance and fault tolerance, although the other two often work themselves in by way of the sort of constraints on the case studies that we're going to look at. Um, You know, all the distributed systems of these problems are, because they have many parts and the parts execute concurrently, because there are multiple computers.
[3:42] You get all the problems that come up with concurrent programming, all the sort of complex interactions and weird timing dependent stuff. And that's part of what makes distributed systems hard. Another thing that makes... distributed systems hard is that Because again you have multiple pieces plus a network. you can have very unexpected failure patterns that is If you have a single computer, it's usually the case that either the computer works or maybe it crashes or suffers a power failure or something. But it pretty much either works or doesn't work.
[4:12] Distributed systems made up of lots of computers, you can have partial failures. That is, some pieces stop working, other pieces continue working. Or maybe the computers are working, but some part of the network is broken or unreliable. um So partial failures is another reason why uh... distributed systems are hard. So this is sort of basic challenges. Bye. Thank you. Thank you.
[4:44] Bye. Subtitles by the Amara.org community Thank you. Thank you. And a final reason why it's hard is that, you know, the original reason to build a distributed system is often to get higher performance. to get 1,000 computers worth of performance, or 1,000 disc arms, worth of performance. But it's actually very tricky to obtain that thousand X speed up with a thousand computers. There's often a lot of roadblocks thrown in your way. Um, Hello.
[5:15] Thank you. Thank you. It often takes a bit of careful design to... make the system actually give you the performance you feel you deserve. So solving these problems, of course, can be all about, you know, addressing these issues. Um, The reason to take the course is because often the problems and the solutions are quite just technically interesting and They're hard problems. For some of these problems, there's pretty good solutions known for other problems They're not such great solutions, no.
[5:46] Distributed systems are used by a lot of real world systems out there, big websites often involved. you know, vast numbers of computers that are, you know, put together as distributed systems. When I first started teaching this course, It was. distributed systems were something of an academic curiosity. You know, people thought, "Oh, you know At a small scale, they were used sometimes, and people felt that, oh, someday they'd be... might be important. but now particularly driven by the rise of giant websites.
[6:17] that have you know, vast amounts of data and entire warehouses full of computers, distributed systems in the last 20 years have gotten to be Very-- seriously important part of computing infrastructure. Okay. This means that there's been a lot of attention paid to them, a lot of problems have been But there's still quite a few unsolved problems. So if you're a graduate student or you're interested in research, There's a lot of problems yet to be solved in distributed systems You could look into his research.
[6:49] And finally, if you like building stuff, This is a good class because it has a lab sequence in which you'll construct some fairly realistic distributed systems focused on performance and fault tolerance. So you've got a lot of practice. building distributed systems and making them work. All right. Let me talk about course structure a bit before I... get started on real technical content. You should be able to find the course website.
[7:20] using Google. And on the course website is the lab assignments course schedule, and also linked to a Piazza where you can post questions, get answers. Um, The course staff, I'm Robert Morris, will be giving the lectures. I also have four TAs. You guys want to stand up and show your faces? Um, Thank you. The TAs are experts at in particular at solving the labs, They'll be holding office hours, so if you have Questions about the labs, you should go to office hours.
[7:54] or you could post questions to Piazza. Um, Thank you. The course has a couple of important components. Um, One is his lectures. There's a paper for almost every lecture. Thank you. Thank you. There's two exams. Thank you. Thank you. Thank you. Bye. There's the labs, programming labs, and Um, there's an optional final project that you can do instead of one of the labs.
[8:27] Thank you. Subtitles by the Amara.org community The lectures will be about big ideas in distributed systems. There will also be a couple of lectures that are more about sort of lab programming stuff. A lot of the lectures will be taken up by case studies, a lot of the way that I sort of try to bring out the um, content of distributed systems is by looking at papers some academic, some written by people in industry, describing real solutions to real problems.
[9:04] Um, These lectures actually would be videotaped. and I'm hoping to post them online so that you can that if you're not here, or you want to review the lectures, you'll be able to look at the videotape lectures. Amen. The papers, again, there's one to read per week. Most of them are research papers. Some of them are classic papers. Like today's paper, which I hope some of you have read on MapReduce. It's an old paper, but it was the beginning of, it spurred an enormous amount of interesting work, both academic and in the real world. So some are classic and some are more recent papers.
[9:38] sort of talking about more up-to-date research, what people are currently worried about. And from the papers, I'll be hoping to tease out what the basic problems are, what ideas people have had that might or might not be useful in solving distributed system problems, Um, We'll be looking at sometimes at implementation details in some of these papers. Because a lot of this has to do with actual construction of of software-based systems. And we're also going to spend a certain amount of time looking at evaluations. people evaluating how fault tolerant their systems by measuring them or people measuring how much performance or whether they got performance improvement at all.
[10:14] Um, So I'm hoping that you'll all read the papers before coming to class. The lectures are... maybe not gonna make as much sense if you haven't already read the lecture, because there's not enough time to both explain all the contents of the paper and have a sort of interesting reflection on what the paper means. online class you really gotta be papers before coming to class and hopefully one of the things you learn in this class is how to read a paper rapidly and efficiently, and skip over the parts that maybe aren't that important sort of focus on teasing out the important ideas.
[10:50] On the website, there's for every link to by the schedule, there's a question. that you should submit an answer for for every paper. I think the answers are due at midnight. And we also ask that you submit a question you have about the paper. Um, to the website in order both to give me something to think about as I'm preparing the lecture and if I have time I'll try to answer at least a few of the questions by email. And the question and the answer for each paper are due midnight. the night before.
[11:21] There's two exams, there's a midterm exam in class, I think on the last class meeting before, uh Spring break. Um, and there's a final exam during final exam week at the end of the semester. The exams are going to focus mostly on papers and the labs. and probably the best way to prepare for them as well as attending lecture and reading the papers. um a good way to prepare for the exams is to look at old exams.
[11:52] 20 years of old exams and solutions. And so you look at those and sort of get a feel for what kind of questions that I like to ask. And indeed, because we read many of the same papers, Inevitably, I ask questions each year that can't help but resemble questions asked in previous years. Thank you. The labs, there's four programming labs. The first one of them is due Friday next week. um They're...
[12:23] Thank you. Lab one is a... Thank you. A simple MapReduce lab. to implement your own version of the paper. in which I'll be discussing in a few minutes. Um, Lab two. involves using a technique called raft in order to get fault, in order to, sort of allow, in theory, allow any system to be made fault tolerant by replicating it and having this raft technique. manage the replication and manage sort of automatic cut over if there's a failed if one of the replicated servers fails.
[12:59] So this is RAF for fault tolerance. Um. Thank you. Subtitles by the Amara.org community *kiss* In lab three, you'll use your raft implementation in order to build a fault tolerant key value server. Thank you. Amen. Thank you. It'll be replicated and fault tolerant. And in lab four, you'll take your replicated key value server and clone it into a number of independent groups. and you'll split the data Um, in your key value storage system across all of these individual replicated groups to get parallel speed up.
[13:37] by running multiple replicated groups in parallel and you'll also be responsible for moving the various chunks of data between different servers as they come and go. without dropping any balls. what's often called a sharded, Bye. key values service. Bye. you Sharding refers to splitting up the data, partitioning the data. among multiple servers in order to get parallel speedup.
[14:10] um Thank you. If you want, instead of doing lab four, You can do a project of your own choice. And the idea here is if you have some idea for a distributed system you know, in the style of... some of the distributed systems we talk about in the class, if you have your own idea that you want to pursue, and you like to build something and measure whether it worked Explore your idea. you can do a project. Um, And so for a project, you'll pick some teammates because we require that projects are done and teams of two or three people.
[14:45] Nam. So like some teammates and send your project idea to us and we'll think about it and say yes or no and maybe give you some advice. um, And then if you go ahead and do, if we say yes, and you want to do a project, you do that instead of Lab 4, and it's do. at the end of the semester. You know, you'll... you should do some, uh, design work and build a real system, and then in the last day of class, you'll demonstrate your system. as well as handing in a short... sort of written report to us about what you built. And I posted on the website some some ideas which might or might not be useful for you to sort of spur thoughts about what projects you might build, but really the best projects are one where you have a good idea, for the project. And the idea is if you want to do a project, you should choose an idea that's sort of in the same vein as the systems that we're talk about in this class.
[15:40] Okay, back to labs. The lab grades that we give you, you hand in your lab code, and we run some tests against it, You're graded based on how many tests you pass. We give you all the tests that we use, so there's no hidden tests. So, If you influence the lab and it reliably passes all the tests, then chances are good, unless there's something funny going on, which there sometimes is. Chances are good that if your code passes all the tests when you run it, it'll pass all the tests when we run it. and you'll get a full score. So hopefully there'll be no mystery about what score you're likely to get.
[16:13] on the labs. Um, Let me warn you that Debugging these labs can be time consuming. Because they're distributed systems, and have a lot of concurrency and communication, Um, sort of strange, difficult to debug errors can crop up. You really ought to start the labs early. Don't leave them. A lot of trouble if you leave the labs to the last moment. You got to start early. If you're-- problems, please come to the TA's office hours.
[16:44] and please feel free to ask questions. about the labs on Piazza. And indeed, I hope if you know the answer that you'll answer people's questions on Piazza as well. All right. Any questions about the mechanics of the course? Yes? Bye. Yeah. Thank you. So the question is, how do these things factor into grade? I forget, but it's all on the...
[17:18] It's on the website under-- something. Um. I think it's the labs are the single most important. component. Thank you. Okay. Thank you. Um... All right, so this is a course about infrastructure. for applications. And so all through this course, there's going to be a sort of split in the way I talk about things between applications, which are sort of other people, the customer, somebody else writes.
[17:49] The applications are going to use the infrastructure that we're thinking about in this course. And so the kinds of infrastructure that tend to come up a lot. Thank you. Thank you. are storage Thank you. communication Thank you. and computation. And we'll talk about systems that provide all three of these kinds of infrastructure.
[18:19] Um, the storage Turns out that storage is going to be the one we focus most on because it's a very well-defined and useful abstraction. And thank you. usually fairly straightforward abstractions. So people know a lot about how to use and build storage systems and how to build sort of replicated fault tolerant high performance distributed implementations of storage. We'll also talk about some of our computation systems, like MapReduce for today.
[18:50] is a computation system. Um, And we will talk about communication some, but mostly from the point as a tool that we need to use to build distributed systems. Like computers have to talk to each other, over a network, you know, maybe you need reliability or something. And so we'll talk a bit about but we're actually mostly... consumers of communication. Um, If you want to learn about communication, systems as or how they work, that's more the topic of 6829. So for storage and computation, A lot of our goal is to be able to discover abstractions, Um, ways of simplifying the interface to these, Storage and computation distributed storage and computation infrastructure so that it's easy to build applications on top of it.
[19:41] And what that really means is that we need to, we'd like to be able to build abstractions that hide the distributed nature of these. of these systems. So the dream, which is rarely fully achieved, but the dream would be to be able to build an interface, That looks to an application as if it's a non distributed storage system just like a file system or something that everybody already knows how to program and has a pretty simple model semantics. We'd love to be able to build interfaces that look and act just like non-distributed, storage and computation systems, but are actually-- you know, vast, extremely high performance, fault tolerant, distributed systems underneath.
[20:23] Um. So we'd love to have abstractions. Thank you. Subtitles by the Amara.org community Thank you. And as you'll see as the course goes on, we sort of, you know only part of the way there. It's rare that you find an abstraction for a distributed version of storage or computation that has simple behavior, behaves just like Um, the non-distributed version of storage that everybody understands. People are getting better at this and...
[20:54] um We're going to try to study the ways in what people have learned about. building such abstractions. Okay. What kind of, what kind of, Topics show up as we're considering these abstractions. The first one, the first topic, general topic that we'll see a lot is in a lot of the systems we look at, have to do with implementation. Thank you. um So, for example, the kind of tools that you see a lot for for ways people learn how to build these systems are things like remote procedure call, whose goal is to mask the fact that we're communicating over an unreliable network.
[21:38] um Another kind of implementation A topic that we'll see a lot is threads. which are a programming technique, that allows us to harness multi-core computers, but maybe more important for this class, Threads are a way of structuring concurrent operations. in a way that's hopefully simplifies the programmer view of those concurrent operations. Um, And because we're gonna use threads a lot, it turns out we're gonna need to also just from an implementation level, spend a certain amount of time thinking about concurrency control, things like locks.
[22:18] Um. you Thank you. And the main place that these implementation ideas will come up in the class, there will touched on in many of the papers, but you're going to come face to face with all of this in a big way in the labs. You need to build distributors. you know, do the programming for distributed systems. These are like a lot of the sort of important tools beyond just sort of ordinary programming, Um, These are some of the critical tools that you'll need to use. I'm back. to build distributed systems.
[22:50] Another big topic that comes up in all the papers we're going to talk about is performance. Thank you. Bye. Thank you. Thank you. Um, Usually the high level goal of building a distributed system is to get what people call scalable speedup. So we're going to look in for scalability. And what I mean by scalability or scalable speedup is that If I have some problem that I'm solving with one computer, and I buy a second computer to help me execute my problem.
[23:30] If I can now solve the problem in half the time, or maybe solve twice as many problem instances. you know, per minute. on two computers as I had on one, then That's an example of scalability. Sort of. two times the computers or resources, um gets me two times the performance or throughput. And this is a huge hammer. If you can build a system that actually has this behavior, Namely that If you increase the number of computers you throw at the problem by some factor, You get that factor, more throughput, more performance.
[24:14] out of the system. That's a huge win. Because You can buy computers with just money. Right? Whereas if in order to get the alternative to this, Um, is that in order to get more performance, you have to pay programmers to restructure your software to get better performance, to make it more efficient, to apply some sort of specialized techniques, better algorithms or something. If you have to pay programmers, the fix your code to be faster.
[24:44] an expensive way to go. We'd love to be able to just, oh, buy a thousand computers instead of 10 computers and get You know, 100 times more throughput. That's fantastic. And so this sort of scalability idea is a huge idea in the backs of people's heads when they're like building things like big websites that run on a you know, building full of computers. If the building full of computers is there to get a sort of corresponding amount of data performance. you have to be careful about the design in order to actually get that performance. Um, So, often the way this looks when we're looking at diagrams or writing diagrams in this course is that That's supposing we're building a website.
[25:26] Ordinarily, you might have a website that you know, has a HTTP server, or let's say it has some has some users. um, running web browsers, And they talked to... you know, web server running Python or PHP or whatever, Sort of web server. And the web server talks to some kind of database. Um, You know, when you have one or two users, have one computer running both, or maybe a computer for the web server and a computer from the database, Maybe all of a sudden you get really popular and you You know, 100 million people sign up for your service.
[26:10] Right? How do you... No. How do you fix your-- you certainly can't-- support millions of people on a single computer, Um, except by extremely careful labor intensive optimization. Um, but you don't have time for. So, Typically the way you're going to speed things up, the first thing you do is buy more web servers and just split the users so that you know, half your users or some fraction of the user go to web server one and the other half You send them to web server two.
[26:40] and because maybe you're building I don't know what, Reddit or something, where all the users need to see the same data ultimately, you have all the web servers talk to the back end. And maybe you can keep on adding web servers for a long time here. Thank you. Um, Thank you. And so this is a way of getting parallel speed up on the web server code. You know, if you're running PHP or Python, maybe it's not too efficient. Um, Uh, as long as each individual web server doesn't put too much load on the database, you can add a lot of web servers.
[27:13] before you run into problems. But. this kind of scalability is rarely infinite, unfortunately. um Certainly not without serious thought. And so what tends to happen with these systems is that at some point after you have 10 or 20 or 100 web servers all talking to the same database, Now all of a sudden the database starts to be a bottleneck. and adding more web servers no longer helps. So it's rare that you get full scalability through sort of infinite numbers of adding infinite numbers of computers.
[27:45] At some point you run out of gas because The place at which you are adding more computers is no longer the bottleneck. by having lots and lots of web servers, we basically move the bottleneck Um, I think it's limiting performance from the web servers to the database. Um, And at this point, actually, you almost certainly have to do a bit of design work. because It's rare that you can take, that there's any straightforward way to take a single database and sort of refactor things. with it for. you can take data sorted in a single database and refactor it so it's split over multiple databases.
[28:21] Um, but it's often a fair amount of work. and Because it's awkward, but many people actually need to do this. Um, We're going to see a lot of examples in this course in which the distributed system people talking about is a storage system. because the authors were running You know, something like a big website that ran out of gas on a single database or storage servers. Um, Anyway, so the Scalability story is we love to build systems that scale this way, but um you know it's hard to make it, or it takes work often, design work, push this idea infinitely far.
[29:04] Amen. Thank you. Thank you. Okay, so. Another big topic that comes up a lot is fault tolerance. Thank you. . If you're building a system with a single computer in it, well, A single computer often can stay up for years. Like, I have servers in my office that have been up for years without crashing. Um, You know, the computer's pretty reliable, the operating system's pretty reliable, Apparently the power in my building is pretty reliable. So it's not uncommon to have single computers that just stay up for amazing amounts of time.
[29:44] However, If you're building systems out of thousands of computers, then even if each computer can be expected to stay up for a year, With a thousand computers that means you're gonna have like about three computer failures per day. And in your set of computers. solving big problems with big distributed systems turns sort of very rare fault tolerance, very rare failure problems, into failure problems that happen just all the time. In a system with a thousand computers, there's almost certainly always something broken. There's always some computer that's either crashed or mysteriously running incorrectly or slowly or doing the wrong thing.
[30:25] Or maybe there's some piece of the network, like with a thousand computers, We got a lot of network cables. and a lot of network switches. You know, there's always some network cable that somebody stepped on and is unreliability, or network cable that fell out, or some network switch whose fan is broken and the switch overheated and failed. There's always some little problem somewhere in your building-sized um, distributed system. So Big scale turns problems from very rare events, you really don't have to worry about that much, into just constant problems.
[30:57] The... failure has to be really, or the response, the masking of failures, the ability to perceive about failures just has to be built into the design. because there's always failures. Um... And, you know, as part of building, you know, convenient abstractions for application programmers, We really need to be able to build infrastructure that as much as possible hides the failures from application programmers or masks them or something. so that every application programmer doesn't have to have a complete complicated story for all the different kinds of failures that can occur.
[31:32] Um, there's a bunch of different notions that you can have about what it means to be fault tolerant. about a little more. but exactly what we mean by that. And we'll see a lot of different flavors, but among the more common ideas you see, one is availability. Thank you. Amen. So, you know, some systems are designed so that under, some certain kinds of failures, not all failures, but certain kinds of failures, the system will keep operating despite the failure while providing undamaged service.
[32:15] Uh, the same kind of service it would have provided, even if there had been no failure. So some systems are available in that sense, So if you build a replicated service that maybe has two copies, You know, if one of the replicas replica servers fail fails, maybe the other server can continue operating They both fail, of course, you can't. You know, you can't promise. Um. availability in that case. So available systems usually say, well, under a certain set of failures, We're going to continue providing service. We're going to be available.
[32:48] More failures than that occur. It won't be available anymore. Thank you. Another kind of fault tolerance you might... you might have or in addition to availability or by itself is recoverability. Thank you. Thank you. Thank you. And what this means is that if something goes wrong, maybe the service will stop working. That is it all. simply stopped responding to requests, Um, and it'll wait for someone to come along and repair whatever went wrong, but after the repair occurs, the system will be able to continue as if nothing bad had gone wrong.
[33:23] So this is sort of a weaker requirement than availability, because here we're not going to do anything while until the failed component has been repaired, But the fact that we can get up get going again without, you know, but... without any loss of correctness, is still a significant requirement. It means, you know, recoverable systems typically need to do things like save their latest date on disk or something where they can get it back you know, after the power comes back up. Um, And even among available systems, In order for a system to be useful in real life, Usually what the way available systems are specced is that They're available until Some number of failures have happened. If too many failures have happened, an available system will stop working or will stop responding at all, But when enough things have been repaired, it'll continue operating. So a good available system will sort of be recoverable as well in the sense that if too many failures occur,
[34:26] um It'll stop answering, but then we'll continue correctly after that. Um, So this is what we love to hear. This is what we'd like to obtain. The biggest hammer, we'll see a number of approaches to solving these problems. I said, sort of Two things that are the most important. tools we have in this department. One is non-volatile storage. so that, you know, something crashed, power failed, or whatever, Maybe there's a building-wide power failure. We can use non-volatile storage like hard drives or flash or solid state drives or something to sort of store a checkpoint or a log the uh state of the system and then when the power comes back up or somebody repairs our power supply, who knows what, we'll be able to read our latest state off the hard drive and continue from there.
[35:22] So. So one tool is sort of non-volatile storage. And the management of non-volatile storage is something that comes up a lot, because non-volatile storage tends to be expensive to update. And so a huge amount of the sort of nitty-gritty of building sort of high performance fault-tolerant systems is in clever ways to avoid having to write the non-volatile storage too much. In the old days and even today you know, what writing non-volatile storage meant was, moving a disc arm and waiting for a disk platter to rotate.
[35:58] both of which are agonizingly slow. on the scale of, you know, 3 gigahertz microprocessors. With things like flash. Life's quite a bit better, but still requires a lot of thought to get good performance out of it. And the other big tool we have for fault tolerance is replication. Thank you. and the management of replicated copies is sort of tricky You know, that sort of problem lurking in any replicated system where we have two servers each with a supposedly identical copy of the system state.
[36:32] Um, The key problem that comes up is always that the two replicas will accidentally drift out of sync and will stop being replicas. And this is just-- you know, at the back of every design that we're gonna see for using replication to get fault tolerance. And Lab 2, and that's what you're all about. management of replicated copies for fault tolerance. Um. As you'll see, it's pretty complex. Um, A final topic, final cross-cutting topic, um is consistency.
[37:11] Thank you. Thank you. So as an example of what I mean by consistency supposing we're building a distributed storage system and it's a key value service so it just supports. two operations, maybe there's a put operation, and you give it a key, and a value and the storage system sort of stashes away the value under as the value for this key, so it maintains just a big table of keys and values. And then there's a get operation. the client sends it a key and the storage service is supposed to respond with the value that's stored for that key.
[37:51] And this is kind of the When I can't think of anything else as an example of a distributed system, I'll whip uh, key value services. And they're very useful, right? They're just sort of the fundamental Um, simple version of a storage system. So. Of course if you're an application programmer. It's helpful if... These two operations kind of have meanings attached to them. You can go look in the manual and the manual says you know, what it what it means, what you'll get back if you call get.
[38:23] right, and sort of what it means for you to call put So it would be great if there was some sort of spec for what they meant, otherwise, like, How can you possibly write an application without a description of what Pudding getters are supposed to do. Um, And this is the topic of consistency. And the reason why it's interesting in distributed systems is that Um. both for performance and for fault tolerant reasons, fault tolerance reason, we often have more than one copy of the data floating around. So, you know, in a non-distributed system where you just have a single server with a single table, there's often, although Not always.
[39:02] There's often relatively no ambiguity about what put and get could possibly mean intuitively. You know, what put means is update the table, and what get means is just, get me the version that's stored in the table, Um, But in a distributed system, where there's more than one copy of the data, due to replication or caching or Who knows what? There may be lots of different versions um of this key value pair floating around. Like if one of the replicas, you know, supposing some client issues a put, and You know, there's-- two copies of the The server, so they both have a key value table.
[39:46] Right, and maybe key one has value 20 on both of them. Thank you. and then some client issues a put. We have a client over here. and it's gonna send a put, it wants to update the value of one to be 21. Maybe it's counting stuff in this Key value server, so it sends a put Thank you. with key one and value 21, it sends it to the first server and it's about to send The same, you know, it wants to update both copies, right?
[40:19] It keeps them in sync. It's about to send this put, but just before it sends the put to the second server, it crashes. power failure, bug in operating system or something. So now the state we're left in, sadly, is that we sent this put and so we've updated one of the two replicas to have value 21, but the other one's still with 20. Now somebody comes along and reads with a "get" And they might get... They want to... Read the value associated with key one, they might get 21 or they might get 20, depending on who they talk to. And even if the rule is you always talk to the top server first, If you're building a fault tolerance system, the actual rule has to be Oh, you talk to the top server first, unless it's failed, in which case you talk to the bottom server.
[41:00] Um, So either way, someday you risk exposing this stale copy of the data. to some future get. It could be that many gets get the updated 21 and then Like next week, all of a sudden, some get yields, you know, a week old copy of the data. So that's not very consistent. Right? in order, but you know, it's, the kind of thing that could happen. Right? and we're not careful. We need to actually write down what the rules are going to be about puts and gets given this danger of due to replication.
[41:37] And it turns out there's many different the definitions you can have of consistency. Um, Many of them are relatively straightforward. Many of them sound like, well, I get yields the you know, value put by the most recently completed put. All right. And so that's usually called strong consistency. It turns out also it's very useful to build systems that have much weaker consistency. That for example, do not guarantee anything like I get it.
[42:11] sees the value written by the most recent put. And the reason, so there's strongly consistent systems - - They usually have some version of get seeing most recent puts. I think you have to... There's a lot of details to work out. There's also weakly consistent, many sort of flavors of weakly consistent systems that do not make any such guarantee. on that. and it may guarantee you well you know, if, someone does a put, then you may not see the put. You may see old values that weren't updated by the put for an unbattered amount of time maybe.
[42:47] Um, And the reason for-- people being very interested in weak consistency schemes. is that Strong consistency, that is, having RIS actually see, be guaranteed to see the most recent right. That's a very expensive spec to implement. um because what it means is almost certainly that you have to, somebody has to do a lot of communication in order to, actually implement. some notion of strong consistency. If you have multiple copies, Um, It means that either the writer or the reader, or maybe both, has to consult every copy. Like in this case, um where maybe a client crashed, left one updated, but not the other.
[43:29] If we wanted to implement strong consistency in maybe a simple way in this system, we'd have readers read both of the copies or if there's more than one copy, all the copies. and use the most recently written value that they find. Um. but that's expensive. That's a lot of chit chat. to read one value. So in order to avoid communication as much as possible, Um. particularly if replicas are far away, people build weak systems that might actually allow the stale read of an old value in this case.
[44:01] Um, although there's often more semantics attached to that to try to make these weak schemes And where this communication problem, you know, strong consistency. requiring expensive communication. Where this really runs you into trouble is that If we're using replication for fault tolerance, then We really want the replicas to have independent failure probability, to have uncorrelated failure. So, for example, putting both of the replicas of our data in the same rack, in the same machine room, is probably a really bad idea.
[44:38] because if someone trips over the power cable to that rack, both of our copies of our data are gonna die. Because they're both attached to the same power cable on the same rack. Um, So in the search for making replicas as independent in failure as possible in order to get decent fault tolerance, People would love to put different replicas as as far apart as possible, like in different cities. or maybe on opposite sides of the continent. So an earthquake that destroys one data center will be extremely unlikely to also destroy The other data center that has the other copy.
[45:13] Um, You know, so we'd love to be able to do that. If you do that, then the other copy is thousands of miles away, And the rate at which light travels means that it may take on the order of milliseconds or tens of milliseconds to communicate to a data center across the continent in order to update the other copy of the data. And so that makes this The communication required for strong consistency, for good consistency, potentially extremely expensive. Like every time you want to do one of these put operations, or maybe a get, depending on how you implement it, You might have to sit there waiting for like 10 or 20 or 30 milliseconds in order to talk to both copies of the data to ensure that they're both updated or both checked.
[45:57] to find the latest copy. Um, And That tremendous expense, right? This is 10 or 20 or 30 milliseconds on machines that, after all, will execute like a billion instructions per second. So we're wasting a lot of potential instructions while we wait. People often build much weaker systems. You're allowed to only update the nearest copy or only consult the nearest copy. Now, I mean, there's a huge sort of amount of academic and real world. research on had us. structure weak consistency guarantee so they're actually useful to applications and how to take advantage of them in order to actually get high performance.
[46:35] Thank you. All right, so that's a lightning preview of the... technical ideas in the course. Any questions about this? before I start talking about MapReduce. All right, I want to switch to MapReduce. That's a sort of detailed case study that's actually going to illustrate most of the ideas. that we've been talking about here. And that produces a system that was Uh...
[47:06] originally designed and built and used by Google, I think the paper dates back to 2004. The problem they were faced with was that they were running huge computations on terabytes and terabytes of data. creating an index of all of the content of the web or analyzing the link structure of the entire web in order to identify the most important pages or the most authoritative pages.
[47:36] And so, you know, the whole web is-- was even in those days tens of terabytes of data. Um... Building an index of the web is basically equivalent to a sort running sort of the entire data. Sort, you know, it's like... reasonably expensive. and to run a sort on the entire content of the web on a single computer I don't know how long it would have taken, but you know, weeks or months or years or something. So, Google at the time was desperate to be able to run giant computations on giant data on thousands of computers.
[48:09] in order that the computations could finish rapidly. It was worth it to them to buy lots of computers so that their engineers wouldn't have to spend a lot of time reading the newspaper or something waiting for their big compute jobs to finish. Um... And so for a while, they had their clever engineers sort of hand write, you know, if you needed to write a web indexer or some sort of link, web link analysis tool. you know, Google bought the computers and they say, "Here, engineers, you know, do--run whatever software you like on these computers."
[48:40] they would laboriously write the sort of one-off manually written software to take whatever problem they were working on and sort of somehow farm it out to a lot of computers and organize that computation and get the data back. Um, If you only hire engineers who are skilled distributed systems experts, Maybe that's OK, although even then, it's probably very wasteful of engineering effort. But they wanted to hire people who were skilled at something else. Um, And not necessarily...
[49:13] Um... engineers who wanted to spend all their time writing distributed system software. So they really needed some kind of framework that would make it easy to just, have their engineers Right? the kind of guts of whatever analysis they wanted to do, like sort algorithm or web index or link analyzer or whatever. Just write the guts of that application and not run, be able to run it on a thousands of computers. um without worrying about the details of how to spread the work over the thousands of computers, how to organize whatever data movement was required, how to cope with the inevitable failures.
[49:49] So they were looking for a framework that would make it easy for non-specialists to be able to write and run giant distributed computations. Um, And so that's what MapReduce is all about. Amen. And the idea is that the programmer just write the application designer, consumer of this distributed computation and just be able to write a simple map function and a simple reduce function, Good. don't know anything about distribution. and the MapReduce framework would take care of everything else.
[50:23] Um, So an abstract view of what MapReduce is up to is It starts by assuming that there's some input and the input is split up into some a whole bunch of different files or chunks in some way. We're imagining that You know, We have input file one, input file two, input file three, Etc. You know, and these inputs are maybe Help.
[50:53] web pages crawled from the web or more likely sort of big files that contain many web--each of which contains many web files. crawl from the web. All right, and the way MapReduce starts is that you define a map function and the MapReduce framework is going to run your map function on each of the input files. you Thank you. And of course you can see here there's some obvious parallelism available.
[51:25] can run the maps in parallel. So each of these map functions only looks at its input and produces output. The output that a map function is required to produce is a list. Takes a file as input and a file as some fraction of the input data, and it produces a list of key value pairs as output. the map function. And so, for example, let's suppose we're writing the simplest possible MapReduce example, a word count MapReduce. job, who's is The goal is to count the number of occurrences of each word.
[51:59] So your map function might emit key value pairs where the key is the word, and the value is just one. So for every word it sees, so this map function will split the input up into words. For every word it sees, it emits that word as the key and one as the value. And then later on we'll count up all those ones. in order to get the final output. Input one has the word. A in it. and the word B in it. And so the output the map is going to produce is Key A value 1, key B value 1.
[52:32] Maybe the second map indication sees a file that has a a b in it and nothing else so it's going to implement output B1, Maybe this third input has an A in it. and a C in it. All right, so we run all these maps on all the input files. And we get this intermediate, what the paper calls intermediate output. which is for every map, a set of key value pairs as output. Then the second stage of the computation is to run the reduces.
[53:04] Thank you. Um, And the idea is that the MapReduce framework collects together all instances from all maps of each keyword. So the MapReduce framework is going to collect together all of the A's you know, from every map, every key value pair whose key was A, it's gonna take Collect them all. and hand them to *knocking* one call of the programmer-defined reduce function. And then it's going to take all the bees and collect them together. Of course, it requires real collection because they were different instances of Key B were produced by different indications of map on different computers. So we're not talking about data movement.
[53:50] So we're going to collect all the B keys and hand them to a different call to reduce that's has all of the B keys. as its arguments, and same with C, all the C's. So this is gonna be the MapReduce framework will arrange for one call to reduce for every key that occurred in Amen. any of the map output. Um, And for our sort of silly word count, example all these reduces have to do or any one of them has to do is just count the number of items passed to it. It doesn't even have to look at the items because it knows that each of them is--the word is responsible for plus one as the value. You don't have to look at those ones, we just count them. So this reduce is gonna produce Um, A and then the count.
[54:43] Of its inputs, this reduce. It's gonna produce the key associated with it and then count of its values, which is also two, Thank you. Thank you. Thank you. So this is what a typical... MapReduce job looks like. the high level. Just for completeness, the--I'm sorry. Well, a little bit of terminology. The whole computation is called the job. any one invocation of map or reduce is called a task. So we have an entire job and it's made up of a bunch of map tasks.
[55:21] and then a bunch of reduced tasks. Um... So an example for this word count, you know, what the map and reduce functions would look like. Thank you. Thank you. Um... Thank you. Thank you. The map function takes a key and a value as arguments. And now we're talking about functions like written in an ordinary programming language like, C++ or Java or who knows what.
[55:52] um So this is just code ordinary people can write. What a map function for word count would do is split the key is the file name, which typically is ignored, we don't really care what the file name was. And the V is the content of "this maps input file." So V just contains all this text. We're going to split V into words. Thank you. Thank you. And then for each word, Thank you.
[56:24] *knocking* *knocking* *knocking* We're just gonna emit. And emit takes two arguments. MIT's call only map can make, and MIT is provided by the MapReduce framework. We get to produce, we hand a MIT a key, which is the word... and a value which is Bye. string one. So that's it for the map function. A word count map function in MapReduce literally could be this simple.
[56:56] Um, So. their sort of promise to make the And, you know, this map function doesn't know anything about distribution or multiple computers the fact we need to move data across the network, or who knows what. This is extremely... straightforward. And the reduce function for... word count um The reduce is called with Remember each reduce is called with sort of all the instances of a given key on the MapReduce framework calls reduce with a the key that it's responsible for and a vector of all the values that the maps produced.
[57:35] associated with that key. The key is the word, the values are all ones, we don't really care about them, we only care about how many there were um And so, reduce has its own emit function that just takes a a value to be emitted as the final output. as the value for the This key. So we're going to admit the length. Thank you. of this array. Thank you. So this is also about as simple as reduced functions are. in MapReduce, namely, Extremely simple.
[58:08] and requiring no knowledge about fault tolerance or Anything else? All right, any questions about... The basic framework, yes. and not... necessary but this might return something. I'm at it. Like France, France, France. using the MapReduce framework that maybe it's reduce with access. that step some number of times. You mean can you feed the output of the reduces sort of-- Yeah.
[58:43] Thank you. Oh, yes. Oh yes, in real life, all right. In real life, it is routine among MapReduce users to you know, define a MapReduce job that took some inputs and produced some outputs and then have a second MapReduce job, you know, if you're doing some very complicated multi-stage analysis. or iterative algorithm. like PageRank, for example, which is the algorithm Google uses to sort of, Estimate how important or influential different web pages are. That's an iterative algorithm is sort of gradually converges on an answer. And if you implement a MapReduce, which I think they originally did, You have to run the MapReduce job multiple times.
[59:28] The output of each one is sort of-- you know, list of web pages, with an updated sort of value or weight or importance for each web page. So it was routine to take this output and then use it as the input to another MapReduce job. We can't do it forever. Oh yeah. it is. all the tears "Well, yeah, you need to sort of set things up the output, you need to-- the reduced function sort of in the knowledge though.
[1:00:00] Meet the produce. data that's in the format or has the information required for the next MapReduce job. I mean, this actually brings up a little bit of a shortcoming in the MapReduce. framework, which is, it's great. if you're Thank you. if the algorithm you need to run is easily expressible as a map followed by this sort of "shuffling of the data by key, followed by a reduce, "and that's it." MapReduce is fantastic for algorithms that can be cast in that form.
[1:00:30] And furthermore, each of the maps has to be completely independent. required to be uh... functional, pure functional functions that just look at their arguments and nothing else. And you know, that's like, it's a restriction. And it turns out that many people want to run much longer pipelines that involve lots and lots of different kinds of processing. And with MapReduce, you have to sort of cobble that together from multiple--and MapReduce. distinct map-produced jobs. and more advanced systems, which we will talk about later in the course, are much better at allowing you to specify the complete pipeline of computations, and they'll do optimization.
[1:01:09] the framework realizes all the stuff you have to do. organize much more complicated efficiently. optimize, much more complicated. Computation. Your question. So in the paper, they distinguish between mappers and map functions. I guess that's the-- like the processes that are running the map. What are the... and welcome the processes that are-- From the programmer's point of view, it's just about map and reduce.
[1:01:43] From our point of view, it's gonna be about the worker processes. and the worker servers that that are part of MapReduce framework, among many other things. call them out from produce functions. So-- um Yeah, from our point of view, we care a lot about how this is organized by the surrounding framework. This is sort of the programmers view with all the distributed stuff stripped out. Um. Yes? Thank you. Sorry, I got to...
[1:02:21] Say it again. Sorry, do you emit locally? Oh, you mean where does the emit data go? Thank you. Okay, so there's two questions. One is... When you call emit, what happens to the data? and the other is where the functions run. So, Thank you. Um. The actual answer is that First, where does the stuff run? There's a number of, say, a thousand servers. Actually, the right thing to look at here is figure one in the paper.
[1:02:59] Um, Though sitting underneath this in the real world, there's some big collection of servers. And Um. We'll call them maybe worker servers or workers. And. There's also a single master server that's organizing the whole computation. And what's going on here is the master server knows that there's some number of input files, you know, 5,000 input files. and it farms out invocations of map to the different workers. So it'll send a message to worker seven saying, please run this map function on such and such an input file.
[1:03:39] And then the worker function which is, you know, part of MapReduce and knows all about MapReduce, well then, Um, Read the file, read the input, whatever, whichever input file. and call this map function with the file name value as arguments. than that worker process. will is what implements emit. And every time the map calls emit, The worker process. will write this data to files on the local disk.
[1:04:11] So what happens to map emits? is they produce files on the map worker's local disk that are accumulating all the keys and values produced by the mapped run on that worker. um So at the end of the math phase, What we're left with is all those worker machines Each of which has the output of some of whatever maps were run on that worker machine. Then the MapReduce workers arrange to move the data to where it's going to be needed for the reduces and since I And a typical big computation This reduced syndication is gonna need all map output that mentioned the key A, but it's gonna turn out, you know, this is a sort of simple example, but probably, In general, every single map invocation will produce lots of keys, including some instances of key A.
[1:05:13] So typically in order, before we can even run this reduce function, the MapReduce framework that is the MapReduce worker running on one of our thousand servers, is going to have to go talk to every single other of the thousand servers and say, "Look, I'm gonna run the reduce for key A, please. Look at the intermediate map output stored in your disk and fish out all of the instances of key A and send them over the network to me. So the reduced worker is going to do that. It's going to fetch from every worker.
[1:05:44] all of the instances of the key that it's responsible for, that the master has told it to be responsible for. And once it's collected all of that data, then it can call reduce. And the reduce function itself calls reduce emit, which is different from the map in it. What reduces emit does is writes the output to A file. in a cluster file service that Google uses. So here's something I haven't mentioned.
[1:06:15] Um, I haven't mentioned where the input lives. and where the output lives. They're both vials. Um, because any piece of input We want the flexibility to be able to read any piece of input on any worker server. That means we need some kind of network file system to store the input data. Um, And so indeed, the paper talks about this thing called GFS, for Google File System.
[1:06:48] Um, And GFS is a cluster file system. And GFS actually runs on exactly the same set of workers worker servers that run MapReduce. And the input GFS just automatically, when you, you know, it's a file system, you can read and write files. it just automatically splits up any big file you store on it. across lots of servers in 64 megabyte chunks. So if you write, if you have 10 terabytes of crawled web page contents. and you just write them to GFS, even as a single big file, GFS will automatically split that vast amount of data up into 64 kilobyte chunks distributed evenly over all of the GFS servers, which is to say, all the servers that Google has available.
[1:07:32] And that's fantastic, that's just what we need. If we then want to run a MapReduce job that takes the entire crawled web as input. the data's already stored in a way that's split up evenly across all the servers. And so that means that um the map workers, you know, we're going to launch If we have a thousand servers, we're going to launch a thousand map workers each reading one one thousandth of the input data. and they're gonna be able to read the data in parallel from a thousand GFS file servers.
[1:08:02] That's getting... A tremendous total. read throughput. you know, the read-through bit of a thousand servers. Yeah. Thank you. running the map. So are you thinking maybe that Google has one set of physical machines that run GFS and a separate set of physical machines that run MapReduce jobs. saying that the May I have two?
[1:08:37] Okay. They're not necessarily . - Right, so the question is, what is this arrow here? actually involve. Um, And the answer to that actually sort of changed over the years as Google's of all of this system. You know, with this in the most general case, if we have big files stored in some big network file system. Like, you know, it's like GFS is a bit like AFS. you might have used on Athena. where you go talk to some collection, your data's split over, big collection of servers, you have to go talk to those servers over the network to retrieve your data.
[1:09:13] Um, In that case, what this arrow might represent is Thank you. the map, the MapReduce worker process has to go off and talk across the network to the correct GFS server or maybe servers. that store its part of the input. and fetch it over the network to the-- MapReduce worker machine in order to pass the map. And that's certainly the most general case. And that was eventually how MapReduce actually worked. In the world of this paper though, And if you did that, that's a lot of network communication.
[1:09:47] You're talking about 10 terabytes of data, and then we have to move 10 terabytes across their data center network, which... you know, data center networks run at gigabits per second, but it's still a lot of time to move. tens of terabytes of data. Thank you. in order to try to, and indeed in the world of this paper in 2004, The most constraining bottleneck in their MapReduce system was network throughput. because they were running on a network, if you read as far as the evaluation section. their network, their network was Um...
[1:10:23] They had thousands of machines, whatever, and they would collect machines, they would plug machines into, you know, each rack of machines into, you know, an Ethernet switch for that rack or something. But then, you know, they all need to talk to each other. But there was a root Ethernet switch. that all of the rack Ethernet switches talk to. And this, and you know, so if you just pick some MapReduce worker and some GFS server chances are at least half the time, the communication between them has to pass through this one root switch. Their root switch had only some amount of total throughput, which I forget, uh... you know some number of gigabits per second.
[1:11:07] um Anyway, I forget the number, but when I did the division, that is divided up the total throughput available in the root switch by the roughly 2,000 servers that they used in the paper's experiments. What I got was that each machine's share of the root switch, or of the total network capacity, was only 50 megabits per second. in their setup. Subtitles by the Amara.org community So 50 megabits per second per machine.
[1:11:39] And that might seem like a lot, 50 megabits, gosh, millions and millions. But it's actually quite small compared to how fast, say, disks run or CPUs run. And so this, with their network, this 50 megabits per second was like a tremendous limit. They really stood on their heads in the design described in the paper to avoid using the network. And they played a bunch of tricks to avoid sending stuff over the network when they possibly could avoid it. One of them was They ran the GFS.
[1:12:12] servers. and the MapReduce workers on the same set of machines. so that they have a thousand machines They'd Run GFS, they implement their GFS service on that thousand machines and run MapReduce on the same thousand machines. And then when the master was splitting up the map work and sort of farming it out to different workers, it would cleverly when it was I'M GOING TO GO TO THE NEXT QUESTION. about to run the map. that was gonna read from input file one, it would figure out from GFS which server actually holds input file one on its local disk.
[1:12:50] And it would-- Send the map. for that input file to the MapReduce software on the same machine. so that by default, this arrow is actually local read from the local disk and did not involve the network. And you know, depending on failures or load or whatever, that it couldn't always do that. But almost all of the maps would be run on the very same machine that stored the data, thus saving them um vast amount of time that they would otherwise have had to wait to move the input data across the network.
[1:13:21] Um, The next trick they played is that map, as I mentioned before, stores its output on the local disk of the machine that you run the map on. So, again, storing the output of the map does not require network communication, at least because the output's stored on the disk. however We know for sure that one way or another, in order to group together all of the way the MapReduce is defined, In order to group together all of the values associated with the given key and pass them single invocation to produce on some machine, this is going to require network communication.
[1:13:59] We're going to--you know, we want to--we need to--we're That's all the A's and give them a single machine to have to be moved across the network. And so this shuffle, this movement of the keys from this kind of, you know, originally stored by Rowe on the same machine that ran the map. We need them essentially to be stored on by column on the machine that's going to be responsible for reduce. This transformation of row storage to essentially column storage is called, the paper calls a shuffle. And it really that required moving every piece of data.
[1:14:30] across the network. from the math that produced it to the reduce that we needed. And that was like the expensive part of the... Um. of the MapReduce. Yeah. Why didn't they? make the something that would run sort of like you would have you Keep your... as you have it already. pick one of the then you would have waited all night. You're right, you can imagine a different definition in which you have a more kind of streaming reduce. I don't know I haven't thought this through. I don't know why. whether that would be feasible or not. Certainly as far as programmer interface, like if the goal, their number one goal really, to be able to make it easy to program.
[1:15:11] by people who just had no idea what was going on in the system. So it may be that this spec, this is really the way reduce functions look. in C++ or something. Like a streaming version of this is now starting to look I don't know how it looked. Probably not this simple. But you know, maybe it could be done that way. And indeed, many modern systems People have gotten a lot more sophisticated. with modern things that are the successors of MapReduce. And they do indeed involve processing streams of data often rather than this very batch approach. It is a batch approach in the sense that we wait until we get all the data and then we process it.
[1:15:55] So first of all, you then have to have a notion of finite inputs, right? And modern systems often do indeed use streams. and... are able to take advantage of some efficiencies due to that. Thank you. but not MapReduce. Thank you. Um... Okay, so this is the point at which the shuffle is where all the network traffic happens, this can actually be a vast amount of data. So if you think about sort, If you're sorting, the output of the sort has the same size as the input to the sort. So that means that if your input is 10 terabytes of data, and you're running a sort, you're moving 10 terabytes of data across the network at this point.
[1:16:37] And your output will also be 10 terabytes and so this is-- quite a lot of data and indeed it is for many map-produced jobs, although not all. There significantly reduce the amount of data at these stages. I'm... Somebody mentioned, oh, what if you want to feed the output of reduce into another MapReduce And indeed, that was often what people wanted to do. And in any case, the output of the reduce might be enormous, like for sort or web-in-mixing. the output that it produces is on 10 terabytes of input. the output of the reduce is again gonna be 10 terabytes. So the output of the reduce was also stored on GFS.
[1:17:10] And the system would-- you know, reduce would just produce these key value pairs, but the, Thank you. MapReduce framework would gather them up and write them into giant files on GFS. Um, And so there was another, round of network communication required to get the output of each reduce to the GFS server needed to store that produce. You might think that they could have... Played the same trick with the output of storing the output on the GFS server that happened to run um the MapReduce worker that ran the reduce.
[1:17:48] And maybe they did do that, but because GFS as well as splitting data for performance, also keeps two or three copies for fault tolerance. That means no matter what, you need to write one copy of the data across a network to a different server. So there's a lot of network communication here, and a bunch here also. Um, And it was this network communication that really limited the... Thruppidimapidus. in 2004. Um, In 2020, Because this network arrangement was such a limiting factor for so many things people wanted to do in data centers.
[1:18:22] Modern data center networks are a lot faster at the root than this was. Um, One typical data center network you might see today actually has many root, instead of a single root switch that everything has to go through, you might have-- many root switches and each rack switch as a connection to each of these sort of replicated root switches and the traffic is split up. among the root switches. So modern data center networks have far more network throughput um And because of that, actually, modern-- I think Google sort of stopped using Macrodeuce a few years ago, but Um, before they stopped using it.
[1:19:01] The modern MapReduce actually no longer try to run the maps. on the same machine as the data was stored on. They were happy to load the data from anywhere because they just assumed the network It was extremely fast. Okay. We're out of time for MapReduce. We have a lab. at the end of next week in which you'll write your own somewhat simplified MapReduce. So, Have fun with that. and see you on Thursday.
Open in the Vidleaf workbench
Search the transcript, select lines, copy quotes with timestamps, translate.
Attribution
"Lecture 1: Introduction" by MIT 6.824: Distributed Systems (https://www.youtube.com/@6.824), licensed under CC BY 3.0 (https://creativecommons.org/licenses/by/3.0/). Source video: https://www.youtube.com/watch?v=cQP8WApzIQQ. This page is a text transcript of the video with paragraph breaks and timestamps added; the creator is not affiliated with and does not endorse Vidleaf.
Are you the creator or a rights holder? Request a correction or removal: copyright@vidleaf.app (see About these pages).
Last updated