TRANSCRIPT · CC BY 3.0

Lecture 6: Fault Tolerance: Raft (1)

MIT 6.824 on Raft, part 1: why earlier systems rely on a single master, how majority voting prevents split brain, and how terms, elections and the log work.

MIT 6.824: Distributed Systems · Published · 1 h 20 min · English · License: CC BY 3.0 · Source: watch on YouTube

Transcript source: automatic speech recognition on Vidleaf (unedited, may contain errors). Paragraph breaks and timestamps added by Vidleaf.

[0:01] All right, let's get started. Mm-hmm. Today, and indeed today and tomorrow, I'm going to talk about Raft. Um, both because I hope it will be helpful for you in implementing the labs. And also because You know, it's just a case study in the details of how to get state machine replication correct. So, by way of introduction to the problem, you may have noticed a pattern in the fault tolerance systems that we've looked at so far.

[0:35] Um, One is that MapReduce replicates computation but The replication is controlled, the whole computation is controlled by a single master. Um, Another example I'd like to draw your attention to is that GFS, replicates data, right, as this primary backup scheme for replicating the actual contents of files, but It relies on a single master. to choose who the primary is for every piece of data. Another example, VMware FT replicates computational right on a primary virtual machine and a backup virtual machine. But in order to figure out what to do next, if one of them seems to fail, it relies on a single test and set server.

[1:20] to help it choose, to help it ensure that exactly one of the primary or the backup takes over if there's some kind of failure. So in all three of these cases, Sure there was a replication system but sort of tucked away in a corner in the replication system there was some scheme where a single entity was required to make a critical decision about. who the primary was in the cases we care about. So, I'm going to start with the first one. a very nice thing about having a single entity decide who's gonna be the primary?

[1:52] is that it can't disagree with itself. There's only one of it, make some decision, that's the decision it made. Um, But the bad thing about having a single entity decide, like, who the primary is, is that it itself has a single point of failure. And so you can view these systems that we've looked at as sort of... pushing the real heart of the fault tolerance machinery into a little corner. That is the single entity that decides who's gonna be the primary. if there's a failure. Now, This whole thing is about how to avoid split brain. The reason why we have to have, have to be extremely careful about making the decision about who should be the primary, if there's a failure, is that otherwise we risk split brain.

[2:35] Um, And just... make this point super clear, I'm going to just remind you. what the problem is and why it's a serious problem. So, supposing for example we're um, We want to build ourselves a replicated test and set server. That is, we're worried about the fact that VMware FT relies on this test and set server to-- choose who the primary is, so let's build a replicated test and set server. And I'm going to do this, it's going to be broken, it's just an illustration for why why it's difficult to get the split brain problem correctly. So, you know, we're gonna imagine we have a network, and maybe two servers, which are supposed to be replicas of our Test and set service connected and you know, maybe two clients that need to know who's the primary right now? or actually, maybe these clients in this case are a The primary and the backup.

[3:30] um, in VMware FT. So, If it's a test and set service, then both these databases, both these servers start out with their state, that is the state of this test and flight bag being zero. the one operation their clients can send is the test and set operation which is supposed to set the of the replicated service to one, so it should set both copies. and then return the old value. So it essentially acts as a kind of simplified lock server. Um, Thank you.

[4:01] OK. The problem situation? What we worry about, split brain arises automatically. when a client can talk to one of the servers but can't talk to the other. So we're imagining either that when clients send a request, they send it to both. Um, I'm just gonna assume that now, it almost doesn't matter. So let's assume that the protocol is that the client's supposed to send ordinarily any request to both servers, And Somehow, we need to think through what the client should do if one of the servers doesn't respond.

[4:32] or what the system should do if one of the servers seems to be unresponsive. Um, Thank you. So let's imagine now the client one can contact server one but not server two. How should the system react? One possibility is that we think, well, you know, gosh, We certainly don't want to just talk to client to server one, because that would leave the second replica inconsistent if we set this value to one, but didn't also set this value to one. So maybe the rule should be that a client is always required to talk to both replicas. to both servers. for any operation and shouldn't be allowed to just talk to one of them.

[5:05] So why is that the wrong answer? Supposing the rule is, oh, in our replicated system, the client's always required to Talk to both replicas. in order to make progress. Thank you. That's not fault tolerant at all. In fact, it's worse. It's worse. than talking to a single server because now the system has a problem if either of these servers is crashed or you can't talk to it. At least with a non-replicated server, you're only depending on one server. But here we have both servers have to be alive. If we require the client to talk to both servers then. Both servers have to be live. So we can't possibly require the client to actually you know, wait for both servers to respond.

[5:48] If we don't have fault tolerance, we need it to be able to proceed. So another obvious answer is that, If the client can't talk to both, well, it just talks to the one it can talk to and figures the other one's dead. So why is that also not the right answer? The troubling scenario is if the other server is actually alive. So suppose the actual problem we're encountering is not that the server crashed, which would be good for us, but the much worse issue that something went wrong with the network cable. And that this client can talk to, client one can talk to server one but not server two.

[6:27] And there's maybe some other client out there that can talk to server two but not server one. So, If we make the rule, that if a client can't talk to both servers, that it's okay in order to be fault tolerant, that it just talked to one, then what's-- inevitably going to happen. said, this cable's going to break, thus Cutting the network in half. I'm going to-- send a test and set request to server one, Server one will set its state to one Return the previous value of 0 to client 1. And so that means client 1 will think it has the lock.

[7:00] and if it's a VMware FT, server will think it can be takeover's primary. But this replica still has zero in it. Right, so now if client two also sends a test and set request, to You know, we-- tries to send them to both, sees that server one appears to be down, follows the rule that says, well, you just send to the one server you can talk to, then it will also think that Um. it acquired, client two will also think that it acquired the lock. And so now, You know, if we were imagining this test and set server was going to be used with VMware FT, we have, you know, Both.

[7:31] uh, replica, both of these VMware, uh, machines I think they could be primary by themselves. Um, without consulting the other server. So that's a complete failure. So with this setup and two servers, It seemed like we had this, we just had to choose, either you wait for both and you're not fault-tolerant, or You wait for just one and you're not correct. And the not correct version is often called split-brain. So. Everybody see this? Thank you.

[8:03] Mm-mm. Well. Um... So this was basically where things stood. until the late 80s. And But people did want to build replicated Systems. the computers that control telephone switches, or the computers that ran banks, you know. Those places are willing to spend a huge amount of money in order to have reliable service, and so they would replicate, they would build replicated systems, and the way they would deal The way that they would-- have replication but try to rule out split brain. It's a couple of techniques. One is They would build a network that could not fail.

[8:44] And so usually what that means, and in fact you guys use networks that essentially cannot fail all the time. the wires inside your laptop you know, connecting the CPU to the DRAM are effectively a network that cannot fail. between the-- between your CPU and DRAM. So, you know, with reasonable assumptions and lots of money, and that, you know, sort of carefully controlled physical situation, like you don't want to have a cable snaking across the floor that somebody can step on, you know, it's got to be carefully physically designed setup. With a network that cannot fail, you can rule out split-brain. It's a bit of an assumption, but with enough money people get quite close to this because If the network cannot fail, that basically means that the client can't talk to server two, that means server two must be down.

[9:31] because it can't have been the network malfunctioning. So that was one way that people sort of built replication systems that didn't suffer from split-brain. Thank you. Um. Another possibility would be to have some human beings sort out the problem, that is, don't automatically do anything, instead have the clients By default, clients always have to wait for you know, both replicas respond or something, I'm never allowed to proceed with just one of them. But you can, you know, call somebody's beeper to go off, some human being goes to the machine room and sort of looks at the two replicas and, either turns one off, to make sure it's definitely dead.

[10:08] or verifies that one of them has indeed crashed, and that the other is alive. And so you're essentially using the human as a as the tiebreaker, and the human is a you know. if they were a computer, it would be a single point of failure themselves. club. So for a long time, people used one or the other of these schemes in order to build replicated systems, and it's not, you know, They can be made to work. The humans don't respond very quickly. And the network that cannot fail is expensive, but it's not not doable. Um, But it turned out that you can actually build automated failover systems Um, that can work correctly in the face of flaky networks, of networks that could fail.

[10:50] that can partition. So this split of the network in half where the two sides operate but can't talk to each other, That's usually called a partition. Um. Thank you. Thank you. And the big insight that people came up with in order to build automated replication systems that don't suffer from split-brain is the idea of a majority vote. And so this is-- Thank you. Thank you. Thank you. a concept that shows up in like every other sentence practically in the Raft paper.

[11:24] It's a sort of fundamental-- um way of proceeding. The first step is to have an odd number of servers instead of an even number of servers. Like, one flaw here is that it's a little bit too symmetric. The two sides of the split here, they just look the same. They run the same software, they're going to do the same thing, and that's not good. But if you have an odd number of servers, Then, It's not symmetric anymore. Right. At least a single network split will be presumably two servers on one side and one server on the other side and they won't be symmetric at all.

[11:57] And that's part of what majority voting schemes are. appealing to. So, basic idea is you have an odd number of servers, in order to make progress of any kind, so in Raft, elect a leader, or cause a log entry to be committed, in order to make any progress at each step, You have to assemble a majority... of the servers, more than half, more than half of all the servers, in order to sort of approve that step, like vote for a leader, accept a new log entry and commit it. Um, So, you know, the most...

[12:29] uh... Straightforward way is to have two or three servers required to do anything. Um, One reason this works, of course, is that If there's a partition, There can't be more than one partition with a majority of the servers in it. That's one way to look at this. A partition can have one server in it, which is not a majority. Or maybe you can have two, but if one partition has two, then the other partition has to have only one server in it. and therefore will never be able to assemble a majority and won't be able to make progress.

[13:03] Um, And just... To be totally clear, When we're talking about a majority, it's always a majority out of all of the servers. not just the live servers. This is a point that confused me for a long time, If you have a system with three servers, and maybe some of them have failed or something, If you need to assemble a majority, it's always two out of three. even if you know that one has failed. The majority is always out of the total number of servers. um, There's a more general formulation of this. Because of majority voting system in which two out of three are required to make progress, it can survive the failure of one server.

[13:40] Right. Any two servers are enough to make progress. If you need to be able to, if you're, you know, you're worried about how reliable your servers are or God. you can build systems that have more servers, and so the more general formulation is If you have two f plus one servers, then you can withstand F failures. Bye. So if it's three, That means f is one. In a system with three servers, you tolerate F servers, F1 failure and still keep going.

[14:16] Thank you. Thank you. Thank you. All right. Often these are called quorum systems because the two out of three is sometimes held a quorum. Thank you. Okay, so one property I've already mentioned about these majority voting systems is that Um, At most one partition can have a and therefore if the network's partitioned, We can't have both halves of the network making progress. Another more subtle thing that's going on here is that Um, If you always need a majority.

[14:49] of the servers to proceed. And you go through a sort of succession of operations in which reach operation somebody assembled a majority like. you know, votes for leaders or Let's say votes for leaders for Raft. At every step, the majority you assemble for that step must contain at least one server that was in the previous majority. any two majorities overlap in at least one server. And it's really that property, more than anything else that Raft is relying on.

[15:22] Um, to avoid split brain. It's the fact that, for example, when you have a leader, a successful leader election, and a leader assembles votes from a majority, its majority is guaranteed to overlap with the previous leader's majority. So, for example, the new leader is guaranteed to know about the term number used by the previous leader. because its majority overlaps with the previous leader's majority and everybody in the previous leaders majority knew about the previous leaders term number. Similarly, anything the previous-- leader could have committed must be present in a majority of the servers in Wrapped, and therefore any new leader's majority must overlap at at least one server.

[16:01] with every committed entry. from the previous leader. This is a big part of Why it is that... Wrapped is correct. Any questions about... the general concept of majority voting system. Yeah? Is it also common to add servers? Like add the processing ? JONATHAN GRUBER: Is it possible to add servers? I think it's possible. It's possible, you know, section something, maybe six in the how to add or change the set of servers.

[16:41] It's possible. You need to do it in a long running system. If you're running your system for five, 10 years, you're going to need to replace the servers after a while. you know, one of them fails permanently, or you upgrade, or you move machine rooms to a different machine room. You really do need to be able to support changing sets of servers. So that's it. It certainly doesn't happen every day, but it's a critical part of this, or long-term maintainability of these systems. Um, And the RAF authors sort of pat themselves on the back There you have it.

[17:12] a scheme that deals with this, which as well they might because it's complex. Thank you. Thank you. All right, so using this idea, in about 1990 or so, there were two systems proposed at about the same time. that realize that You could use this majority voting system to kind of get around the apparent impossibility of avoiding split brain. by using, basically by using three servers instead of two, and taking majority votes.

[17:44] Um, in the one of these very early systems was called Paxos The Raft paper talks about this a lot. Um, And another of these very early systems was called view stamp replication. Ciao. abbreviated as VSR for view stamp replication. And even though Paxos is by far the more widely known system in this department, Raft is actually closer to design in design to view step by few stamp replications. which was invented by people at MIT.

[18:15] Um, And so there's sort of a long, many decade history of these systems and they only really came to the forefront and started being used a lot in deployed big distributed systems about 15 years ago. a good 15 years after they were originally invented. Okay, so... So let me talk about Rath now. Raft takes the form of a library intended to be included in some service application.

[18:50] And so if you have a replicated service, each of the replicas in the service is going to be some application code which receives RPCs or something, plus a raft library. And the raft libraries cooperate with each other Um, to maintain replication. So, a sort of software overview of a single-raft replica is that at the top we can think of the replica as having the application code. So it might be for lab three, a key value server. So maybe we have some key value server.

[19:22] And it has state, the application has state that raft is helping it manage, replicated state. And for a key value server, it's going to be a table of keys and values. Um... *knocking* The next layer down is a raft layer. So the key value server is going to make function calls into RAPT, and they're going to chit chat back and forth a little bit. raft keeps a little bit of state. You can see it in figure two. And for our purposes, really the most critical piece of state is that raft has a log of operations.

[20:01] Thank you. Okay. Thank you. Amen. And in a system with three replicas, we're actually going to have three servers that have exactly the same identical structure and hopefully the very same data sitting in. sitting at both layers. Thank you. Amen. Thank you. Thank you. Thank you. Thank you. Thank you.

[20:31] Thank you. Right. Outside of this, there's going to be clients. And the game is that So we have client one, client two, a whole bunch of clients. The clients don't really know, the clients are, you know, External code that needs to be able to use the service. and The hope is the clients won't really need to be aware that they're talking to a replicated service. That to the clients it'll look almost like it's just one server and they talk to one server. And so the clients actually send client requests to The key to the application layer.

[21:05] of the current leader, the replica that's the current leader in Raft. And so these are going to be application level requests for a database for a key value server, these might be put and get requests. put takes a key and a value and updates the table and get ask the service to get the current key, current value corresponding to Some key. So this has nothing much to do with Raft. It's just client-server interaction for whatever service we're building.

[21:39] Um, But once one of these commands gets sent from the request, gets sent from the client to the server, what actually happens is on a non-replicated server, the Application code would like execute this request and say update the table in response to a put. but not in a RAPT-replicated service. Instead, if assuming the client sends a request to the leader, what really happens is the application layer. simply sends the request, the client's request, down into the RAFT layer to say, "Look, here's a request.

[22:10] Please get it committed. into the replicated log and tell me when you're done. And so at this point, the rafts chit chat with each other. Um. until all the replicas or a majority of the replicas get this new operation into their logs So it is replicated, and then when it's And the leader knows that all of the replicas of a copy of this only then is wrapped send a notification up, back up to the key value layer saying, "Aha, that operation you sent me a minute ago, It's been now committed into all the replicas And so it's safely replicated and at this point it's okay to execute that operation.

[22:53] It's a raft. The client sends a request to the key value layer, QDialyLayer does not execute it yet. So we're not sure. because it hasn't been replicated. Only when it's in the logs of all the replicas. Bye. Then Raph notifies the leader, and now the leader actually executes the operation, which corresponds to... you know, for a put, updating the key value table for a get, reading the correct value out of the table, and then finally sends the reply back to the client.

[23:26] So that's an ordinary operation of the Thank you. . It's submitted if it's in a majority. And again, the reason why I can't be all is that if we want to build a fault tolerance system, It has to be able to make progress even if Some of the servers have failed. So-- Um, Yeah. It's committed when it's in a majority. Thank you. Yes. So before the leader executes the request markets, devalue properties.

[24:00] Um... that have been replicated. Thank you. Wow. - I'm sorry. Thank you. leaders Yeah, I got that. And so. In addition, when the operations finally committed, each of the replicas sends the operation up. Each of the wraths Library layer sends the operation up to the local application layer, and the local application layer applies that operation to its state. And so they all... So hopefully all the replicas seem the same stream of operations.

[24:31] they show up in these up calls in the same order, they get applied to the state in the same order, and you know, assuming the operations are deterministic, which they better be, Um, The state... of the replicated state will evolve. in identically on all the replicas. So typically this table is what the paper is talking about when it, Talks about state. Thank you. a different way of viewing this interaction and one that'll sort of notation that will come up a lot in this course is that Um, A sort of time diagram, I'll draw you a time diagram of how the messages work.

[25:12] So let's imagine we have a client, and Server one is the the leader, and we also have server two. Thank you. Server three, and time flows downward on this diagram. We imagine the client sending the original request. to server one Thank you. After that, server one's wrapped layer. Um, sends an append entry, it's RPC. Um, to each of the two replicas. Thank you. This is just an ordinary, let's say, put request.

[25:45] This is, these are PEND entries requests. Thanks. The server's now waiting for replies. from the servers, from the other replicas, as soon as replies from a majority, a ride back, including the leader itself. So in a system with only three replicas, the leader only has to wait for one other replica. to respond positively to an appendentary as soon as it assembles. positive responses from a majority. Um, the Leader executes the command, figures out what the, Answer is like forget.

[26:20] um and sends the reply back to the... And meanwhile, of course, you know, if S2 is actually alive, it'll, you know, Sendback gets response too, but we're not waiting for it. Thank you. It's useful to know in figure two. Thank you. All right, everybody see this? Thank you. This is the sort of ordinary operation of the system. No failures here. Oh, okay. Gosh, yeah. I like...

[26:53] I left out-- important step. At this point, the leader knows, oh, I got a majority of put it in their log, I can go ahead and execute it and reply yes to the client because it's Server two doesn't know anything yet. It just knows, well, you know, I got this request from the, but I don't know if it's committed yet. Depends on, for example, whether my reply got back to the leader. For all Server 2 knows, its reply was dropped by the network, Maybe the leader never heard the reply and never decided to commit this request. There's actually another stage, once the, Server.

[27:24] realizes that a request is committed, it then needs to tell the other replicas that fact. And Amen. So there's... Uh... There's an extra message here. Exactly what that message is depends a little bit on what else is going on. at least in RAF, there's not an explicit commit message. Instead, the information is piggybacked inside the next append entries that the leader sends out. The next append entries RPC it sends out for whatever reason.

[27:58] Like there's a, commit, leader commit or something field in that RPC. and The next time the leader needs to have to send a heartbeat or something, needs to send out a new client request because some different Client. request something, it'll send out the Um, New, higher things. leader commit value and at that point the, uh, Um, yeah. replicas will execute the operation and apply it to their state. Yes.

[28:29] Thank you. Thank you. Oh, yes. So this is a, This is a protocol that has quite a bit of chit-chat in it. Um, and it's not, Super fat. Indeed, you know, yeah, the client sends in a request, the request has to get to the server, the server talks to at least, you know, another... It sends out multiple messages, has to wait for the responses, sends something back. So there's a bunch of message round-trip times kind of.

[29:02] embedded here. would have made sense to the second part of the faculty. Thank you. Yeah, so... If... So, This is up to you as the implementer actually exactly when the leader, sends out the updated commit index. If-- Client requests a um, come back only very occasionally. then the leader may want to send out a heartbeat or send out a special append entries message.

[29:34] Um, If client requests come quite frequently, then it doesn't matter, If they come, you know, if there's a thousand arrive per second, then, geez, there'll be another one along very soon. And so you can piggyback. So without generating an extra message, which is somewhat expensive. you can get the information out on the next message you are going to send anyway. Um, In fact, I don't think the, time at which the replicas execute the request - It's critical because nobody's waiting for it, at least if there's no failures. If there's no failures, the replicas executing the request isn't really on the critical path.

[30:12] Like the client isn't waiting for them, the client's only waiting for the leader to execute. So it may not be that important. um It may not affect client perceived latency, sort of exactly how this gets staged. Um. Bye. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. All right. One question you should ask is why is the system so focused on logs? What are the logs doing?

[30:51] And it's sort of worth trying to come up with explicit answers to that. One answer to why the system is totally focused on logs. is that The log is, the kind of mechanism by which the leader orders operations. It's vital for these replicated state machines that all the replicas apply not just the same client operations to their start, but the same operations in the same order. But they all have to apply these operations coming from the clients in the same order. And the log, among many other things, is part of the machinery by which the leader assigns an order to the incoming client operations.

[31:32] you know, 10 clients send operations to the leader. At the same time, the leader has to pick Pick an order, make sure everybody, all the replicas obey that order. And the log is The fact that the log has numbered slots. is part of how the leader expresses the order it's chosen. Um, Thank you. Another use of the log is that between this point and this point, Server 3 has received an operation that it is not yet sure is committed. and it cannot execute it yet. It has to put this operation aside somewhere.

[32:06] until the increment to the leader commit value comes in. And so, another thing that the log is doing is that, on the followers, the log is the place where the follower sort of sets aside operations that are still tentative. that have arrived but are not yet known to be committed. And they may have to be thrown away, as we'll see. Um, So that's another use. On the sort of dual of that use on the leader side is that the leader needs to remember operations in its log because it may need to retransmit them.

[32:38] to followers. If some follower is offline, maybe it's something briefly happened to its network connection or something, misses some messages. The leader needs to be able to resend blog messages that any followers miss. And so the leader needs a place where I can set aside copies of messages of client requests, even ones that it's already executed. in order to be able to resend them to the client. I mean, we send them to replicas that missed miss that operation. And a final reason for all of them to keep the log is that at least in the world of figure two, If a server crashes and restarts and wants to rejoin um You really want a server that crashes to in fact restart and rejoin.

[33:20] the wrapped cluster, otherwise you're now operating with only two out of three servers and you can't survive any more failures. We need to reincorporate failed and rebooted servers, and the log is sort of where, what a server, a rebooted server uses. The log persisted to its disk. Because one of the rules is that Each Rath server needs to write its log to its disk, where it will still be after it crashes and restarts. That log is what the server uses, it replays the, operations in that log from the beginning to sort of create its state as of when it crashed.

[33:52] and then it carries on from there. So the log is also used as part of the persistence plan. as a sequence of commands to rebuild the state. Yes. So I'm wondering what scenarios would arise where the workers are trying to follow up with it. Uh... Well, ultimately, Okay, so the question is, Suppose the leader is capable of executing a thousand client commands a second and the followers are only capable of executing a hundred client commands per second. That sort of sustained rate at full speed.

[34:32] um The, so one thing to note is that the, the The replicas, the followers, acknowledge commands before they execute them. So the rate at which they acknowledge and accumulate stuff in their logs is not limited. So, you know, maybe they can acknowledge it, a thousand requests per second. If they do that forever, then they will build up. unbounded size logs because their execution rate falls will fall an unbounded amount behind the rate at which the leader has given the messages, sort of under the rules of our game.

[35:06] And so what that means is it will eventually run out of memory. At some point, after they have a billion, after they fall a billion log entries behind, they'll just like, they're called a memory allocator. for space for a new log entry and it will fail. Um, So yeah, and RAF doesn't, um RAP doesn't have the flow controls that's required to cope with this. So I think... in a real system, you would actually need probably piggybacked and doesn't need to be real time, but you probably need some kind of um, additional communication here that says, "Oh, here's how far I've gotten in execution.

[35:47] So that the leader can say, well, I'm too many thousands of requests ahead of the point at which the followers have executed. Yeah, so I think there's probably in a production system that you're trying to push to the absolute max. you might well need an extra message to throttle the leader if it got too far ahead. Thank you. Okay.

[36:33] - So the question is, If one of these servers crashes, It has this log that it persisted to dis, 'cause that's one of the rules of figure two. um So the server will be able to be just logged back from disk, but of course, That server doesn't know how far it got in executing the log. And also it doesn't know, at least when it first reboots, by the rules of Figure 2. It doesn't even know how much the log is committed. So the first answer to your question is that immediately after a restart, after a server crashes and restarts and reads this log, it is not allowed to do anything with the log.

[37:09] because it does not know how far the system has committed in its log. Maybe this log has a thousand uncommitted entries and zero committed entries for all So... If the leader dies, of course, that doesn't help either. Let's suppose they've all crashed. This is getting ahead of... It's getting a bit ahead of me, but we'll suppose they've all crashed, and so all they have is the state that was marked as non-volatile in figure two, which includes the log and maybe the latest term.

[37:43] Um, And so they don't know. So if there's a crash, if they all crash and they all restart, None of them knows. initially how far they had executed before the crash. So what happens is, they do leader election, one of them gets picked as a leader. And that leader, if you sort of track through what figure two says about how append entries are supposed to work. the leader will actually figure out as a byproduct of sending out the first heartbeat really.

[38:14] Um, It'll figure out what the latest point is basically that that Amen. all of the, that a majority of the replicas agree on their law. Because that's the commit point. Another way of looking at it is that once you choose a leader, through the append entries mechanism, the leader forces all of the other replicas to have identical logs to the leader. And at that point, plus...

[38:45] A little bit of extra, the paper explains. At that point, since the leader knows that it's forced all the replicas to have a have logs that are identical to it. Then it knows that, oh, all the replicas must also have a-- There must be a majority of of replicas with- that all those log entries in the logs which are now identical must also be committed because they're held on a majority of replicas. Um, And at that point, the leader-- append entries code described in figure two for the leader will increment the leader's commit point and and everybody can now execute the entire log from the beginning.

[39:22] and recreate their state from scratch. possibly extremely laboriously. Um... So that's what figure two says. It's obviously this re-executing from scratch is not very attractive, but That's what the basic protocol does. And we'll see tomorrow. version of this is more efficient, uses checkpoints and We'll talk about it tomorrow. Thank you. Okay, so this was a sequence in sort of ordinary, non-failure operation.

[39:53] Um, um... Another thing I want to briefly mention is what this interface looks like. You've probably all seen a little bit of it. due to working on the labs, but roughly speaking, if you have, let's say, this key value layer, with its state. and the raft layer underneath it, there's On each replica there's really two main pieces. of the interface between them. There's this... Um, method by which the key value layer can relay. If a client sends in a request, the key value layer has to give it to Rath and say, "Please, you know, fit this request into the log somewhere, and that's the start function.

[40:34] Um. that you'll see in raft.go. And really it just takes one argument, the client command. Thank you. The key value layer is saying, "Please, I got this command. Stick it into the log and tell me when it's committed." And the other piece of the, interface is that by and by, The raf layer will notify the key value layer that, aha, that operation that you sent to me in a start command a while ago, which may well not be the most recent start. A hundred client commands could come in and cause calls to start before any of them are committed.

[41:09] by and by Um, this-- upward communication is takes the form of a message on a Go channel that the Raph library sends on and the key value layer. is supposed to read from. Um. So there's this apply, Thank you. Thank you. It's called the apply channel, and on it you send an apply message. Thank you. I don't know. Thank you. Thank you. Thank you. this start and of course you need the Key value layer needs to be able to match up messages it receives on the Apply channel with calls to start that it made.

[41:47] Um, And so the start command actually returns enough information for that matchup to happen. It returns the index the start functions basically returns the index in the log where if this command is committed, which it might not be, it'll be committed at this index. Um, And I think it also returns the current term, and some other stuff we don't care about very much. And then this apply message is going to, contain the index. both the Command.

[42:20] Um. Thank you. Thank you. Thank you. Thank you. Thank you. and the index. And all the replicas will get these apply messages, so they'll all know, "Oh, I should apply this command." um figure out what this command means and apply it to my local state. And they also get the index. The index is really only useful on the leader so it can, figure out. what client requests we're talking about. Thank you. I'm curious about the following. The client is sending their request to the back, rock leader any place.

[42:59] at which-- clients by making a get-go. Right. about that. OK. OK. Um, I think I answered a slightly different question. Let's suppose the client sends any request in, Let's say it's a put or a get. It could be a put or a get. It doesn't really matter. I'll take it to get. Um, the point at which the So the client sends in again and waits for a response. The point at which the leader will send a response at all is after the leader knows that command is committed.

[43:40] So this is going to be a sort of get reply. Bye. *knocking* Bye. So the client doesn't see anything back. I mean it, and so, um, That means in terms of the actual software stack, That means that the key value, the RPC will arrive, the key value layer will call the start function. The start function will return to the key value layer But-- The key value layer will not yet reply to the client because it does not know if it's Actually, it hasn't executed the client's request now. It doesn't even know if it ever will.

[44:14] because it's not sure if the request is going to be committed. And the situation in which it may not be committed is if the key value layer, you know, gets a request, calls start, and immediately after start returns, it crashes. It certainly hasn't sent out its apply, pen messages, or whatever. Nothing's been committed yet. Um, So the game is start returns. time passes, the relevant, apply message corresponding to that client request appears to the key value server on the apply channel and only then and that causes the key value server to execute the request And-- Send a reply.

[44:52] Thank you. Thank you. Thank you. Thank you. Thank you. And that's like-- all this is It's very important when it It doesn't really matter if everything goes well. But if there's a failure, we're now-- At the point where we start worrying about failures, we'll be extremely interested in, if there was a failure, what did the client see? Um, Thank you. All right. And so one thing that has come up is... that's, um, all of you should be familiar with is that, at least initially, one interesting thing about the logs is that they may not be identical.

[45:30] There are a whole bunch of situations in which, at least for brief periods of time, the ends of the different... replica's logs may diverge, like for example, if a leader starts to send out a round of append messages, but crashes before it's able to send all of them out. That'll mean that some of the replicas that got the append message will append you know, that new log entry, and the ones that didn't get that pin message as RPC, won't have appended them. So it's easy to see that the logs are going to diverge sometimes. The good news is that, the The way Raft works actually ends up forcing the logs to be identical after a while.

[46:07] 'cause there may be transient differences, but in the long run, all the logs will sort of be modified by the leader until the leader ensures they're all identical and only then Um. Okay. I think the next, there's really two big topics. to talk about here for Raft. One is how leader election works, which is Lab 2. And the other is, how the leader deals with the different replicas logs, particularly after failure.

[46:38] So first I want to talk about leader election. Amen. Thank you. Question to ask is how come the system even has a leader? Why do we need a leader? that Part of the answer is you do not need a leader to build a system like this. It is possible to build. an agreement system by which a cluster of servers agrees on, you know, the sequence of entries in a log without having any kind of designated leader. And indeed, The original Paxos system, which the paper refers to, original Paxos did not have a leader.

[47:11] Um, So it's possible. The reason why Raft has a leader is basically that Um, There's probably a lot of reasons, but one of the foremost reasons is that you can build a more efficient in the common case in which all The servers don't fail. It's possible to build a more efficient system if you have a leader. Because with a designated leader, everybody knows who the leader is, You can basically get agreement on requests with one round of messages per request. Whereas leaderless systems have more of the flavor of, well, you need a first round to kind of agree on a temporary leader, and then a second round to actually send out the request.

[47:49] So, It's probably the case that use of a leader speeds up the system by a factor of two. Thank you. And it also makes it sort of easier to think about what's going on. Thank you. Um, Raft goes through a sequence of leaders, And It uses these term numbers in order to sort of disambiguate which leader we're talking about. It turns out the followers don't really need to know the identity of the leader. They really just need to know what the current term number is.

[48:20] Each term has at most one leader. That's it. critical property You know, for every term, there might be no leader during that term, or there might be one leader, but there cannot be two leaders. during the same term. Every term has at most one leader. Um, Thank you. Thank you. How do the leaders get created in the first place? Every... wrapped server. keeps this election timer, which is just a It's basically just a time that it has recorded that says, well, if that time occurs, I'm going to do something.

[48:55] um And something that it does is that if an entire leader election period expires, without the server having heard any message from the current leader then the server sort of assumes probably that the current leader is dead and starts an election. So we have this election timer. And if it expires, start an election. Thank you. .

[49:26] Thank you. Um, And what it means to start an election is basically that you increment the term the candidate, the server that's decided it's going to be a candidate and sort of force a new election, first increments this term because, you know, wants her to be a new leader, namely itself, And a term can't have more than one leader, so we've got to start a new term. in order to have a new leader. And then it sends out these request votes, RPCs.

[49:57] Thank you. I'm going to send out a full round of request votes. And you may only have to send out n minus one request votes, because one of the rules is that a new candidate always votes for itself. in the election. Um, So. One thing to note about this is that It's not quite the case that if a leader didn't fail, we won't have an election. or if the leader does fail, then we will have an election assuming any other server is up, because someday the other server's election timers will go off.

[50:28] But if the leader didn't fail, we might still, unfortunately, get an election. So if the network is slow or drops a few heartbeats or something, we may end up having election timers go off, and even though there was a perfectly good leader, we may nevertheless have a new election. So we have to sort of keep that in mind when we're thinking about correctness. Um, And what that in turn means is that if there's a new election, it could easily be the case that the old leader is still hanging around and still thinks it's the leader. Like, if there's a network partition, for example, And the old leader is still alive and well in a minority partition.

[51:02] The majority partition may run an election, and indeed a successful election, and choose a new leader. totally unknown to the previous leader. So we also have to worry about, you know, oh, what's that previous leader going to do since it does not know there was a new election? Thank you. Very much. Thank you. Thank you.

[51:43] So the question is, Are there, can there be pathological cases in which, for example, One way network communication can prevent the system from making progress. I believe the answer is yes. Certainly. So for example, if the current leader, If its network somehow half fails in a way the current leader can send out heartbeats but can't receive any client requests. then the heartbeats that it sends out, which are delivered because it's outgoing network connection works. It's outgoing heartbeats will suppress any other server from starting an election.

[52:20] But the fact that it's incoming network wire apparently is broken will prevent it from hearing and executing any client commands. It was absolutely the case that RAPT is not proof against all um sort of all crazy network problems that can come up. I believe the ones I've thought about, I believe are fixable in the sense that the We could solve this one by having a sort of requiring a two-way heartbeat in which if the leader sends out heartbeats but in which Followers are required to reply in some way to heartbeats, like as they are already required to reply.

[52:58] But if the leader stops seeing replies to its heartbeats, Then after some amount of time in which the season was replaced, the leader decides to step down. I feel like that's specific issue. can be fixed. and many others can too, but I, but, Yeah. you know You're absolutely right. Very strange things can happen to networks, including some that the protocol is not prepared for. Thank you. Thank you. Thank you.

[53:29] Okay, so we got these meter elections. We need to ensure that there is at most, at most, one leader per term. How does RAF do that? Well, RAF requires, in order to be elected for a term, RAF requires a candidate to get yes votes from a majority of the servers. The servers and each server will only cast one yes vote per term. So in any given term, you know, it basically means that in any given term, each server votes only once for only one candidate, Um, You can't have two candidates both get a majority of votes because Everybody votes only once.

[54:06] Um, So the majority's-- The majority rule causes there to be at most one winning candidate. Um, and So then we get at most one. candidate elected per turn. Um... And in addition, critically, the majority rule means that you can get elected even if some servers have crashed. If a minority of servers are crashed or unavailable, have network problems, we can still elect a leader. If more than half have crashed or are not available or are in another partition or something, then Actually, the system will just sit there trying again and again to elect a leader.

[54:46] and never elect one. if it cannot in fact, if they're not a majority of live servers. um Thank you. If an election succeeds, Everybody, it'd be great if everybody learned about it. 7 minutes or so. We need to ask ourselves how do all the parties learn what happened? The server that wins an election, assuming it doesn't crash, The server that wins election we'll actually see a majority or positive votes for its request vote from a majority of the other servers. The candidate that's running the election that wins it, the candidate that wins the election will actually know directly, I got a majority of votes.

[55:23] But nobody else directly knows who the winner was or whether anybody won. Um, So the way that the candidate informs other servers is that heartbeat. the rules in Figure 2 say, oh, if you win an election, you're immediately required to send out independentries. to all the other servers. Now, dependentries that heartbeat dependent entries doesn't explicitly say, I won the election, you know, I'm a leader for term 23. Um, It's a little more subtle than that. the way the information is communicated is that Um, No one is allowed to send out and append entries unless they're a leader for that term.

[56:02] So the fact that I'm a, you know, I'm a, Server and I saw, oh, there's an election for term 19. And then by and by, I send in append entries whose term is 19, That tells me that somebody, I don't know who, but somebody won the election. Um, so that's how the other servers know, is they were receiving append entries for that term. And that append entries also has the effect of, resetting everybody's election time. timer. So as long as the leader is up and it sends out heartbeat messages or append entries at least, you know, at the rate it's supposed to, Every time a server receives an append entry, it'll reset its selection timer and sort of suppress anybody from being a new candidate.

[56:46] So as long as everything's functioning, the repeated heart beats will prevent any further elections. Of course, if the network fails or packets are dropped, there may nevertheless be an But if all goes well, we're sort of unlikely to get an election. Um, This scheme could fail in the sense, well, it can't fail in the sense of electing. two leaders for a term, but it can fail in the sense of electing zero leaders for a term. Um, The sort of morning way it may fail is that if too many servers are dead or unavailable or have bad network connections. So if you can't assemble a majority, you can't be elected, nothing happens.

[57:22] Um, The more interesting way in which an election can fail is if Everybody's up. You know, there's no failures, no packets are dropped. But-- two leaders become candidate close together enough in time that they split the vote between them. or say three liters. Um, So supposing we have three leaders. Supposing we have a three replica system Um, All their election timers go off at the same time.

[57:53] Every server votes for itself. And then when each of them receives a request vote from another server, well it's already cast its vote for itself and so it says no. So that means that all three of the servers each get one vote each, nobody gets a majority, and nobody's elected. Um, So then-- Their election timers will go off again because The election timer is only reset if it gets an append entries, but there's no leader, They'll all have their election timers go off again. And if we're unlucky, they'll all go off at the same time. They'll all vote for themselves. Nobody will get a majority. uh So.

[58:24] So clearly, I'm sure you're all aware at this point, there's more to this story. And the way raft D-- makes the possibility of split votes unlikely, but not impossible, is by randomizing these election timers. So the way to think of it And the randomization, well, The way to think of it is that, so we have some timeline, I'm going to, um, draw events on. There's some point at which everybody received the last append entries.

[58:56] And then maybe the server died. Let's just assume the server sent out a last heartbeat and then died. Well, all of the followers have-- or reset their election timers when they received, at the same time, because they probably all received the append entries at the same time. They all reset their election timers. for some point in the future, to go off at some point in the future. they chose different random times in the future which then were gonna go off. So I suppose the dead leader is server one, So now server two and server three at this point, set their election timers for a random point in the future. Let's say server two set their, um, I'd like some timer to go off here.

[59:38] and server three set. It's election timer to go off there. and The crucial point about this picture is that assuming they pick different random numbers, One of them is first and the other one is second. Right? That's what's going on here. And the one that's first, assuming This gap is big enough. the one that's First, this election timer will go off first before the other one's election timer. as long as we're not unlucky, it'll have time to send out a full round of Vote requests.

[1:00:10] and get answers from everybody where everybody's alive before the second election. timer goes off from any other server. Um, So. Does everybody see? how the randomization desynchronizes these candidates. Unfortunately, there's a bit of art in setting the constants for these election timers. There's some sort of competing requirements you might want to fulfill. Um, So one obvious requirement is that the election timer has to be at least as long as the expected interval between heartbeats.

[1:00:48] This is pretty obvious that the leader sends out heartbeats every 100 milliseconds you better make sure, you know, there's no point in having the election timer anybody's election timer ever go off for 100 milliseconds. because then it will go off before we could even have expected a new append entry. So the lower limit is certainly the lower limit is one heartbeat interval. In fact, because the network may drop packets. You probably want to have the minimum election timer value be a couple of times the heartbeat interval. So for 100 millisecond heartbeats, you probably want to have the very shortest possible election timer be say, 300 milliseconds.

[1:01:25] three times the heartbeat interval. Um, So that's the sort of minimum is the... Um, See if the heartbeats are this frequent, you want the minimum to be a couple of times that or here. Um, so what about the maximum? You know, you're gonna presumably randomize uniformly over some range of times. You know, where should we set the-- kind of maximum. Thank you. that we're randomizing over. And there's a couple of...

[1:01:55] considerations here in a real system. um, you know, this maximum time, uh... affects how quickly the system can recover from failure because Remember, from the time at which the server fails, until the first election timer goes off, The whole system is frozen. There's no leader. You know, the client's requests are being thrown away because there's no leader. and we're not assigning a new leader even though you know Presumably these other servers are up.

[1:02:26] So the bigger we choose this maximum, the longer delay we're imposing on clients before recovery occurs. You know, whether that's important, depends on sort of how high performance we need this to be, and how often we think there will be, Failure. If failure's happening once a year, then who cares? Um, If we're expecting failures frequently, we may care very much. how long it takes to recover. Okay, so that's one consideration. The other consideration is that this gap That is the expected gap in time between the first timer going off and the second timer going off.

[1:03:03] Um, This gap really, in order to be useful, has to be longer than the time it takes for the candidate to assemble votes from everybody. That is... longer than the expected round trip time, the amount of time it takes to send an RPC and get the response. Yeah, maybe it takes 10 milliseconds. to send in RPC and get a response, get a response from, um, all the other servers. And if that's the case, we need to make maximum release long enough that there's pretty likely to be 10 milliseconds difference between the smallest random number and the next smallest random number.

[1:03:35] Um, Thank you. And for you, Um, the test code will get upset if you, uh, um Thank you. if you don't recover from a leader failure in a couple seconds. And so, just pragmatically, you need to tune this maximum down so that It's highly likely that. You'll be able to complete a leader election within a few seconds. That's not a very tight constraint.

[1:04:09] Any questions about the election timeouts? One tiny point is that, um, You want to... Choose new random timeouts. every time there's, every time you Every time I know-- resets its election timer. That is, Don't choose a random number when... server is first created and then reuse that same number over and over again. Because you make an unlucky choice that is You choose the same.

[1:04:39] one server happens by ill chance to choose the same random number as another server. That means that you're gonna have split votes over and over again. forever. So that's why you want to almost certainly choose a different, a new, a fresh random for the election timeout value every time You reset the timer. Thank you. All right, so, a final issue about leader election. Suppose we are in this situation where the old leader's partitioned.

[1:05:10] You know, the network cable is broken, and the old leader is sort of out there with a couple clients. and a minority of servers. There's a majority in the other half of the network, and the majority in the new half of the network elects a new leader. What about the old leader? Why won't the old leader... Thank you. cause incorrect execution. Thank you. Maybe it's Yeah, right there.

[1:05:56] That's how that's gonna... - I'm glad to. Bye. and that's all done. God. Thank you. Yeah. two potential problems. Or one sort of non-problem is that if there's a leader off in another partition and it doesn't have a majority, then the next time a client sends it a request, Um, that leader that, you know, an abarctuation with a minority, yeah, it'll send out append entries, but Because it's in the minority partition, it won't be able to get responses back from a majority of the servers, including itself. And so it will never...

[1:06:34] commit the operation, it'll never execute it. It'll never respond to the client. saying that it executed it either. And so that means that, yeah, an old server off in a different partition People may, clients may send a request, but they'll never get responses. um, So no client will be fooled into thinking that the, that old server executed anything for it. The other sort of more tricky issue, which actually I'll talk about in a few minutes, is the possibility that before a server fails, it sends out append entries to a subset.

[1:07:14] of the servers. and then crashes before making a commit decision. And that's a very interesting question, which um We'll probably spend a good 45 minutes talking about that. And so, actually, before I turn to the back, topic in general. more questions about leader election. Thank you. OK. um Okay, so how about the contents of the logs and how in particular how a newly elected leader possibly picking up the pieces after an awkward crash of the previous leader, How does a newly elected leader sort out the possibly divergent logs on the different replicas?

[1:08:01] in order to restore sort of consistent state in the system. Um, All right, so the first question is, What can things, this whole topic is really only interesting after a server crashes, right? If the server stays up, then relatively few things can go wrong. If we have a server that's up and has a majority, you know, during the period of time when it's up and has a majority, Um, just tells the followers what the log should look like.

[1:08:36] And the followers are not allowed to disagree. They're required to accept. They just do by the rules of figure two if they've been more or less keeping up. You know, they just take whatever the leader sends them, and append entries, and append it to the log, and obey commit messages, and execute. Hardly anything to go wrong. The things that go wrong in rap go wrong when the old leader crashes sort of midway through you know, sending out messages or a new leader crashes, you know, sort of just after it's been elected, Um... before it's done anything very useful.

[1:09:08] So, one thing we're very interested in is what can the logs look like after some sequence of crashes? Um... Okay. Here's an example. So then we have three servers. Thank you. Thank you. Um, And the way I'm gonna draw out these diagrams, 'cause we're gonna be looking at a lot of, sort of situations where the logs look like this, and we're going to be wondering, is that possible, and what happens if they do look like that? So, my notation's going to be, I'm going to write out log entries for, um, Each of the servers sort of aligned to indicate slots.

[1:09:46] corresponding slots in the log. And the values I'm gonna write here are the term numbers rather than client operations. I'm going to, you know, this is slot one, this is slot two. Um, Everybody saw a command from term three in slot one, Server 2 and Server 3 saw command from also from term three, In the second slot, the server one has nothing there at all. Um, And so the question for this, like the very first question is, can this arrive?

[1:10:18] Could this... set up Arise and how? Oh. How? Thank you. Yeah, so . Thank you. So.

[1:11:03] So maybe server three was the leader for, just repeating what you said, maybe server three is the leader for term three. It got a command that sent out to everybody. Everybody received the dependent as the log. And then I got a server three, Got a second request from a client. and Maybe it sent it to all three servers, but the message got lost on the way to server one, or maybe server one was down at the time or something, and so only server two Well, the leader always appends new commands to its log before it sends out append entries. Um, and maybe the append entry RPC only got to server two. So this situation, you know, It's like the simplest situation in which actually the logs are now different.

[1:11:38] And we know how it could possibly... arise. And so if server three, which is the leader, should crash now, you know, the next server they're going to need to make sure server one Well, First of all, If server 3 crashes, or we get an election and some other leader is chosen, You know, two things have to happen. The new leader... It's got to recognize that This. Um, a command could have committed. It's not allowed to throw it away. and it needs to make sure server one fills in this blank here with indeed this very same command that everybody else had in that slot.

[1:12:15] Um, Thank you. All right, so after a crash, somebody, you know, a server suppose, another way this can come up is server three might have sent out the append entries to server two, but then crashed before sending the append entries to server three. So if we're electing a new leader, it could be because We got a crash. before the message was sent. All right, here's another scenario to think about. Um, Three servers again. Thank you. Thank you. Thank you. Um, Now I'm going to number the slots in the lawn.

[1:12:47] So we can refer to them. Um, We've got slot 10, 11, 12. Thank you. 13. Again. That's... Thank you. same setup except Now we have in slot 12, we have Server 2 as a command from term four and server three has a term command from term five. Thank you. Thank you. So-- Before we analyze these to figure out what would happen, what would a server do if it saw this, we need to ask Could this even occur?

[1:13:25] Sometimes the answer to the question, oh geez, what would happen if this configuration arose? Sometimes the answer is, It cannot arise, so we do not have to worry about it. um So the question is, Could this arise and how? Yeah. say no, because we have a lot of . Okay. Thank you. Thank you. Um, All right.

[1:14:00] And any-- All right. Thank you. The first situation happens. Amen. watching. friends. - Now, Troy. Bye. Roger. Thank you. What? Uh... Right, sensei.

[1:14:32] guys. Thank you. - Oh my God. All right. back up. Yeah. How far? So let me just repeat that. Brief. we know this configuration can arise, and so the way we can then get a four and a five here is, Let's suppose in the next leader election, server 2 is elected leader.

[1:15:09] Now for term four. Select the leader. You get a request from a client, it appends it to its own log and crashes. So now we have this. Right, we need a new election because the leader just crashed. Now, in this election, Thank you. And so now we have to ask whether who could be elected, or we have to keep in the back of our heads, oh gosh, who could be elected? Um, so we're gonna claim server three could be elected. The reason why it could be elected is because it only needs request vote responses from a majority. That majority is server one and server three.

[1:15:41] You know, there's no problem, no conflict between these two logs. So server three can be elected for term five, get a request from a client. appended to its own log and crash. And that's how you get this. uh, this configuration. So you need to be able to-- Bye. to work through these things. in order to get to the stage of saying, yes, this could happen, and therefore, Raft must do something sensible, as opposed to, it cannot happen. Because some things can't happen.

[1:16:12] Um, Thank you. All right, so... Um, So what can happen now? We know this can occur. Hopefully we can convince ourselves that Raft actually does something sensible. Now, As for the range of things, before we talk about what Raph would actually do. We need to have some sense of what would be an acceptable outcome.

[1:16:47] Right? and just eyeballing this, Um, We know that The command in slot 10, since it's known by all All the replicas, it could have been committed. so we cannot throw it away. Similarly, the command in slot 11, since it's in a majority of the replicas, it could, for all we know, have been committed, so we can't throw it away. The command in slot 12, however, neither of them could possibly have been committed. So we're entitled. We don't know, we haven't Looked at what Raft will actually do, but Raft is entitled to drop both of the even though it is not entitled to drop in either of the commands in the 10 or 11.

[1:17:26] Um, This is an entitled drop. It's not required. to drop either one of them. And we know it certainly must drop one, at least one, because You have to have identical log contents in the end. This could have been committed. the We can't tell by looking at the logs. exactly how far the leader got before crashing. So one possibility, is that for this command or even this command, One possibility is that a leader send out the append messages with a new command and then immediately crashed.

[1:18:04] So it never got any response back because it crashed. So the old leader did not know if it was committed. And if it didn't get a response back, That means it didn't execute it and it didn't send out but, you know, it didn't send out the incremented commit index. And so maybe the replicas didn't execute it either. So it's actually possible that this wasn't committed. So, Even though Raft doesn't know, Um, It could be legal for Raft Thank you.

[1:18:36] If RAP knew more than it does know, It might be legal. to drop this log entry, because it might not have been committed. but because on the evidence, There's no way to disprove it was committed based on this evidence. It could have been committed. and Raft can't prove it wasn't. So it must treat it as committed. Right, because the leader might have received it, might have crashed just after receiving the append entry replies. and replying to the client. So just looking at this, we can't rule out the possibility either possibility.

[1:19:11] that the leader responded to the client, in which case we cannot throw away this entry because a client knows about it. Or the possibility the leader never did, and yeah, we could. you know So we have to assume that. It was committed. Question? . um, Uh...

[1:19:46] No, there's no--I mean, in this situation, maybe the server crashed before getting the All right, let's continue this. on Thursday.

Open in the Vidleaf workbench

Search the transcript, select lines, copy quotes with timestamps, translate.

Open in the workbench →

Attribution

"Lecture 6: Fault Tolerance: Raft (1)" by MIT 6.824: Distributed Systems (https://www.youtube.com/@6.824), licensed under CC BY 3.0 (https://creativecommons.org/licenses/by/3.0/). Source video: https://www.youtube.com/watch?v=64Zp3tzNbpE. This page is a text transcript of the video with paragraph breaks and timestamps added; the creator is not affiliated with and does not endorse Vidleaf.

Are you the creator or a rights holder? Request a correction or removal: copyright@vidleaf.app (see About these pages).

Last updated