Machine learning - introduction
Nando de Freitas's first machine learning lecture: why programs must learn from data, shown with vision, recommendation, web-scale text and animal perception.
Transcript source: automatic speech recognition on Vidleaf (unedited, may contain errors). Paragraph breaks and timestamps added by Vidleaf.
[0:00] I'm going to play this and then just follow the instructions here at the top. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. OK, so which changes did you-- perceived color of Taller of the doors.
[0:36] At this rate. Pardon? I'm not going to be screaming. Oh, yeah, the sign, the participates sign change. Thank you. What else changed? And this is for everyone, 340. Thank you. The what? The railing disappeared. Thank you. So in particular-- There was-- this red link has disappeared. the sign Is that friend? Thank you.
[1:09] The building on the right. How many people think that this building on the right is completely different than the initial Go ahead. There's 340 people. Okay. How many people think there was a dog here in front? Up this door. There's three, four people. There wasn't. I was just tricking you. There's a person there. This was the original image. So as you can see, the building wasn't pink. It was a different building.
[1:43] There were windows on top instead of doors. The rail was complete. there were people there was a petition signed And-- that all change in front of your eyes. and most of you were completely oblivious to most of these changes. I'm going to restart it and play it one more time. just to illustrate the point that even if you know that these changes are taking place. you will yet again not be able to see them all.
[2:18] You will see some, but you won't see them all. And it's because your eye does not see the world. You don't have the world in your head. Your head is too small. the world cannot fill it in. So what you do is your eye moves. and gazes at the world about three times second and then you compose with those cases an image of what the world is about. And most of what you see is actually imagined. And imagination plays a big part in this course. This course, to a large extent, is about imagination.
[2:50] If you can imagine the world. that means that you have learned. what the world is about. The videos also illustrate an important thing And it's that-- If you wanted to program a machine that sees We are confronting a serious problem here. Because we think we know how we see and what we see. And yet I've just proven to you that that notion is actually wrong.
[3:23] that we actually don't know how we see. see and now if we do not know how we see How can we code? and or Python or MATLAB. and tell a machine to see. comes extremely hard because we can't introspect And if we can't introspect, it's very hard to replicate a brain. Thank you. All right. That-- his minute introduction to the topic of today.
[3:57] which will be just your-- first lecture. And in today's lecture, I will-- tell you a bit about what is machine learning. And where you can use machine learning? I'll talk about the sort of the big data revolution, why machine learning has all of a sudden become and attested with the fact that we now have access to gazillions of data. computing power to process that data. Thank you. I'll talk about--about the--about the--about the--about the--about the--about the Part of the course-- something that's going to be a Big part of the course, neural systems, and how we draw inspiration from biological systems in order to build learning machines.
[4:40] And finally, I'll just-- touch on a few applications. which hopefully will start already giving you some motivation for what to do for your project. Okay. So how do we tell machines how to see? things So one of the big problems is that the world is not easy to describe. If you wanted to tell a machine how to recognize giraffe, that's going to be very hard because there's giraffe of all shapes. It might not be as bad as these giraffe here, which have been exaggerated.
[5:17] But it's very hard to tell a computer what a giraffe is. because they come in all sorts of shapes or in different backgrounds You could have a giraffe in Africa, but all of a sudden, you get a giraffe in the snow. Calgary soon. And it might be an albino giraffe, in which case, color would then help you. Um. So how do we tell--how do we code? what Azure Office is. It turns out to be extremely hard. or And even simpler problem, how do we code-- what a face How do we tell a machine There's a face in that image.
[5:54] And that's sort of-- something that you take for granted because-- all your iPhones or phones have face detector. Your cameras these days come with FaceTime. detector. You use Facebook, it has face detectors. It knows where the faces are. in the images. time. How many of you know how to build a face detector? There's just three, four students their hands up. You've seen random forests. Don't put fuel out. Maybe we'll do that in this course as a homework exercise.
[6:28] Build a Facebook. It's very easy with machine learning technology. But if you try to do it by coding it, it's extremely hard. So one anecdote that I always tell people is that, Once I was doing-- consulting for a car manufacturer in Japan and when I sat I was doing a visit there to the factory And I sat at the car. And it immediately knew where my face is. So there's a camera. and intent so that they know what I'm doing, what I'm looking. and whether I'm falling asleep.
[6:59] warfare. tea goods on. And so, but then, A friend of mine, Jesse, who's now a professor in Waterloo, And Jesse's tall and blonde. And this thing could not find his face. And the reason is because there were only Japanese engineers in this. And they coded the faces essentially dark eyebrows, and eyes, and so on. So the system, the hand-corder system was extremely brittle. So it's like a system coded to recognize your op.
[7:32] and optimal conditions. but impossible to recognize you're out when running away from a predator. Bye. The machine learning solution is-- Collect a large collection of face images. These will be from all sorts of different types of people. and populations and so on. and then come up with a very simple algorithm that will allow you to detect face. So you can't tell a machine-- How to?
[8:06] You can't write a program that will tell machines exactly what a face is, but you can code a program that will... tell explain to machine these are faces These are not faces. And then let the machine automatically learn what it is that differentiates a face from a non-face. And that's the machine learning approach. It's based on examples as opposed to we telling the rules. We let the machines figure out which are the things, which are the features, we call them features or patterns, or statistics that make a face different than any other image.
[8:47] So the first part is to collect that huge collection of images of faces. You then collect a big data set of known faces And-- The result is this. That's how your face-- detectors are built. Um. I will show you how to do that in the course with a technique called random forests. and so you can have nice consumer electronics and so on. Um. Of course, it's not just for faces. You could do this for anything. any object. So you might take lots of images of pedestrians, and train a machine to recognize pedestrians.
[9:24] Now, why would I want a machine to recognize pedestrians? Thank you. What do you mean? automated cars. Automated cars are not far. They're already approved in California. Um, and, um... Soon. uh... will hopefully have automated cars on the street. Um... I am trying to remember a figure twenty or forty Either 20 or 40.
[9:56] I've got my power set to confuse. Either 20,000 or 40,000 people die every year in the United on car accident. mostly pedestrian. To put this into the importance of this project, in context, imagine the population of UBC dissipated every year. That's the cost. to not driving properly and relying on humans Try it. In India, when I checked a couple years ago when I was visiting, It was 113,000 people.
[10:31] that career which may or may not be bad depending on how you count. Now, um... Actually, it is bad regardless of how you count. 113,000 rises. Huge number of people. Thank you. So-- Same idea. *coughs* You teach a machine, here's a bunch of images, in this case images of cars, And then again, commercial systems these days, like Mobile already are able to detect flames, detect cars and so on.
[11:06] So-- But the key to all these huge innovations, be able to have an automatic-- because without these innovations, it wouldn't be possible at all to have an automatic car. Because how could you deploy a car that can't see a and understand what a pedestrian is in front of it. that would be like murder. Um... the key to it was to come up with this realization that we could teach machines. what they take. So this course, indeed, will be about how to write the programs that let machines grab data And.
[11:42] decide what a face is and and/or what a pedestrian is, and so on. Um. Now, there's other learning machines out there, and those learning machines are humans. Humans are very good. In fact, not just humans, all animals are very good learners. Um. Humans are particularly good. Um. and I'm about to now test. how good humans are at learning with this example. So, I have highlighted here, I have a bunch of images.
[12:15] of objects. which three of them I'm calling tufas. Okay, you're not labeled them by putting a red box around them. So now the question to you guys is how many of you think this one here that I'm pointing with the arrow is a tufa? Okay. the majority. How many of you think this is a two-fly? How about this one? That's a Skeptics.
[12:48] How about this? That a tearful? Just one? So go to an interesting sample, this one here. Thank you. Okay, no, not as many people. Um... This guy here. - This guy? Dear people, Yeah. OK, let's pick someone who likes Let's start with that.
[13:21] Why is that a tufa? God bless you. stem, and it's kind of got roots. and OK, stem, root, and a head. That's what makes it. to do fat. Uh... Who else would like to comment on why they said something is a twofer? or not. - Looking at the dog, the There's a long frown. There is a loose.
[13:52] just sort of Red Bulls. Thank you. Come on. It's hard to describe, but somehow you know. Because you-- very confident to put your hand up or kept your hands down. Even though you couldn't describe it, say that that was a tufa or not a tufa. See? Do you know what a two-finger is? Thank you. most people Anyone else?
[14:23] Thank you. Thank you. So... This illustrates several things. One is how quickly you learn what a two-fay is. You've never seen a two-foot people. maybe in your data field, But you very quickly learn what a 2P is. Thank you. from very few examples. Um... Not all of you have the same opinion on what a truffe is, that's why I asked several questions. So it's somewhat subjective. But there is nonetheless strong agreement Um... And then... There are features that make a twofer be a twofer, whether it's a long stem or the legs head Um...
[14:59] *cough* But by and large, we're not thinking even of those features when we extract them consciously. It's down subconscious. A lot of the feature extraction interpretation of the world is done through the subconscious. Subconscious is very powerful. perhaps more than the conscious. Um. We would like to design machines that are capable of doing this. that can see a few examples of data And just from a few examples, can very quickly learn new concepts.
[15:32] Machines are not capable of making these induction. LEAPS. as Josh and Yvonne calls it. So machine learning is about Being able to extract the features or the patterns or the statistics. The thing that makes us tell what a twofer is, And once you have those features, And you might not do this consciously. We humans don't do this consciously and that's why It's hard. called to try to code things by hand, and instead we have to rely on machines.
[16:07] is to extract the patterns that make a two-fold. The patterns that make a face or a pedestrian And then once you have those, to make a prediction. Where a prediction could be like a decision. deciding whether something is a face or not a face. whether something is a tooth or not a face. or a prediction of a different nature as in, is this thing going to eat me? Should I run away? Is this too dangerous or not? Thank you.
[16:39] There's other types of predictive tasks that are very popular right now, so I'm just going some of them quickly. Um. forecasting. Um... And this I will do as an example in the next class. We're quite often viewing it-- forecast how much energy demand. your building needs, because based on that you can adjust Um... you know, how much energy you buy. And-- you know, potentially safe. Um... Spend the ends of the dollars. for the economy.
[17:10] Thank you. You might want to forecast sales. If you work for Lululemon, you're going to forecast how many people will go for the pink That's a-- at Taskware. Machine learning is actually very useful. Certainly the GAAP and so on stores across North America very-- series of other machine learning. Thank you. You might want--you might want to--you might want to you might have rated a few movies And so you would want a machine to automatically rate the rest of them to protect what you would have how many stars you would have assigned to any other movie so that they can build a better recommendation system.
[17:48] Um... you might want to use machine learning once you teach machine learning what is an HIV virus you want the machine to be able to then detect any departure from this thing that it's monitoring from being a virus. In other words, you want machines to detect mutations. Now those are very important because the moment you detect a mutation you you can choose a different treatment uh... treat the virus.
[18:21] Um. The reason why using your credit cards is so easy these days is because there's very good credit scoring systems. So machine learning classifies that quickly. will determine whether you're-- uh... whether you pass the test or not, whether you're credit worthy or not. So a lot of our economy already-- relies on the usage of very large scale classical Cash in. as well as medical you know, diagnosis.
[18:51] on like noces of mammograms and so on is something that machines eventually do a lot better than um Humans. Ranking, another task where machine learning is very useful. So if you use Google image search for example, quite often you do a search. and then you click on an image. When you do that, you're basically telling Google, Hey Google, for this word, This is the image that matches it.
[19:21] And this is how Google Image Search got better and better over the time because we were--we All humans were labeling the images by clicking. Often you do a search for a machine learning term. And at the top, the top 50 images shown, there's these pictures of hot models. And that's because that's where people click when they search for support vector machines or whatever technique they want. So there's some failures to the algorithm, which is what makes what made me realize what it is that they were doing.
[19:54] Thank you. uh... And of course-- Um... In this huge flow of information that we live in today, being able to summarize information is also important. And already there's lots of tools that will try to condense information for you, whether it's news or any other type of information and make it available. One of these tools is-- um, was Zite, and that was built by two students who take took this course many years ago, my class and Eric Berchu.
[20:26] They, together with some venture-- capitals here in BC eventually built this company that was sold to CNN. If you go to CNN website, the algorithms that power CNN were developed here. Okay, so... and finally Um. Often talking about AI, I think machine learning is the real AI. certainly a huge component of AI.
[20:58] or there is needed in order to build intelligent machines, but it's certainly a huge component And great progress has been made in the last decade. Thanks so much, Imran. Thank you. Now, when do we need machine learning? When our expertise is absent. If you want to send a robot to Mars, Um... Remote control is not going to work, because it takes too long to transmit the signal. So you need a robot that can understand the environment, and it can repair itself, and so on. Thank you. Um... From the video, it was clear that we can't understand how we see or recognize.
[21:34] of things, or even we don't even understand how We-- see concepts from sounds. And so-- lack of ability to introspect motivates the use of learning Um... It's also useful when you're dealing with problems that keep changing. Like if you're trying to give someone good recommendations for, I don't know, kitties or dogs or whatever it is that people recommend these days on the web.
[22:06] Um. People's interest change with time. And so if you relied on curators, Um... that would be--would require a lot of courage. So it's important to Um. to automate to have automatic systems so that As people's interests change with time, you can still provide them with the content they need. Um. Connected to that is-- this problem that often everyone is different. If I wanted to recommend news, I wouldn't have to build a news recommendation system for each different person.
[22:41] because each person That means that if I have a million customers, I will need, say, to the-- instead of having a million curators, what I want to try to have is a million automatic recommending contents. Um... And of course-- Um. for problems where we humans can't reason. And when you work for a search engine, for a company doing web analytics, then it soon becomes clear that you can't rely on rules of-- you can't rely on people to do things The scale of the problem is so large, you do really have to of relying on machine learning.
[23:31] Now, data. Data has-- increased a lot over the last few years. In Um, as-- Lon, Halaby, Peter Norbeck, and Fernando Pereira, researchers at Google, from some of the leading researchers on language at Google. Um. point out in the paper. in 2009. In 1960, seven Um. One million words was considered a large data set. In 2006, I don't even know what that number is.
[24:03] Um-- Trillion. That's it. Thank you. uh... Thank you. Word dataset is what was regarded a large dataset set in academia. So there's been huge orders of magnitude. on the data sets that you guys are capable of actually managing in your personal conversations. And then if you go to Amazon EC2 and so on, it gets even more fun. What were the success stories of language.
[24:37] I'm out. over the last few years. technological success. Watson. Pardon? What's that? Watson? Jeffrey Lee Chan. Cute. success story. Success stories as-- And that they've actually created revenue for a lot of people and solved people's problems. Google Translate. ERIC SCHMIDT: But Watson is a good one. Google Translate. That's a brilliant one. Because... Um.
[25:07] My wife can be at Aritzia doing her job, and she needs textiles from India, and she needs someone to weave those textiles in China, And she's sending emails to all these people in India China. who might be speaking Hindi, who might be speaking Mandarin, and they're all able to communicate because they're all using Google Translate. So it completely revolutionizes e-commerce. because now the world becomes a smaller place. You don't even need to all speak the same language in order to communicate.
[25:40] Occasionally, there's some problems with that. That's pretty much how people do things. even in small companies in Vancouver, this is how--this is how--this is how we're going corporations are able to engaging global business. with their product. What's another example? Apple marketed it. heavily. suit Speech recognition. So, it's picture recognition theory and So-so. Many particles.
[26:12] We're in Canada. But-- Googles. speech recognition works extremely well. I don't know if you've tried it recently. If speech recognition on mobile phone. The beautiful thing as well is that every time you get it right or wrong, You provide feedback. for Google. get better. So I'm assuming that speech recognition is one of those solved. Um. How-- you Thank you.
[26:45] combined teacher of the year Cheers. It's very strange. by my books. Pardon? Something like real time translation. into another language. And... Oh, yeah. You can even do both of those. So eventually we'll be able to speak if you like. Yeah, like science fiction. Where we can all... Maybe I'll tap talk. Father. I wouldn't be surprised.
[27:17] In terms of speech recognition using deep nets, Um. something I'm going to come to later in the course. There's two teams doing it. The first team that started it was actually in soft. And then Google got started on it. And so they're using slightly the same-- sort of infrastructure. about that. common thing among all these? is that there's lots of data. There's lots of data in two languages. And there's-- lots of-- in other words, match data. So this is the text in French and this is the translation.
[27:55] um And the same for speech recognition. People created this huge data. sets. Moreover, the data sets keep getting larger by the minute because where companies like Google and Microsoft and so on have been brilliant is that they give you a service and if you don't like the translation They allow you to correct it. And now because of that, you're always providing data for these companies to get better. at providing their products. So they're harnessing the crowds.
[28:26] Thank you. uh... Now, for a lot of these applications, you do need a lot of data. A lot of data is essential. If you're going to do speech recognition, Really, the stumbling block is not even knowing how to build the technology. Um... but it's actually having the data. to train technology. Data is the commodity. Yeah. Here's an example to illustrate this point. by Alyosha Efers at CMU. So he had initially an image that he didn't quite like because it had a house So what he wanted was-- could I remove this house?
[29:05] So you can obviously remove it. And then the question is, If you have a collection of other images from the web, Can you use this other collection to fill this in? You know. in a way that will look better than this. If-- he found that if he used a few images, it didn't work. But then when he increased this, She was able to have the machines automatically composing new images That's-- still look very plausible. For this to happen, you really need a lot of data.
[29:42] Now, coming back to the Google guys, they say, In their paper-- that they've solved the problem of how to get people to contribute data. And-- Certainly Facebook and Google have done a great job at this. Um... Dave? They claim to have solved the technological problem of aggregating and indexing the data. They have done this to some extent, and Google certainly has done a good job. There's still a lot to be done.
[30:14] But most importantly is the fact that We still have not interpreted the content. So Here's an example, citation matching. Now, when you do a search in Google, you essentially use an inverted index and you look for that term in the search and you check if that exact word appears in the web page and then you bring back those web pages. And then there's some ranking that gets done using an algorithm which--if you do the Python exercise in the Python website, that's the--that will be a way of learning to learn how to code a Google search engine.
[31:04] Now... However, the problem with doing search this way is it's very much based on popularity. Google is a popularity machine. Now, a counter example is citation matching. So citation matching is the problem of trying to decide whether two citations. So in scientific papers, you always cite other people's work at the end. So this is something you will do when you write your report for your project.
[31:38] Deciding whether two citations are the same is actually proved to be extremely hard. And that's because people often misspell names. Often people write the volume or not, include the pages, or have the wrong pages. And what tends to happen a lot with academics is if I cite a work and I didn't get the page number right, And then someone else. wants to cite my work, but my work really goes back to the other paper. They haven't read the other paper, but they just cite it nonetheless. And what they do is they copy the citation.
[32:14] And also you're just a couple of hours behind the deadline, you have to submit a paper that's just cut and paste with a bunch of citations. You'll see what I mean when it comes to handing the project at the end. And so the problem with this is that if something is wrong, it can very easily propagate. Now you have 898 citations with the wrong page number and 22 with the right page number. If you just go with popularity, you will pick those. So you have to do something else. You have to actually do an intervention to be able to solve that problem.
[32:54] Let me show you another example of this. um This is the issue of information extraction. And I'd love to see a few projects on this topic in the course. So this information extraction engine is by Oren Ezioni and Dan Weld, who are professors at the University of Washington in Seattle.
[33:26] And so this allows you to type things like Obama. Married. And now we're starting to do something a bit more interesting, something with semantic value, not just a query, like, you know, not just-- putting in like Obama married and retrieving a bunch of pages that say Obama married but you actually want to know who married Obama who or who did he marry and if you do a search And they do this by extracting triplets of noun, verb, noun from the web.
[34:08] And so that is what you get and the numbers I'm imagining is the frequency. That's the web. The web has truth, but the web is also full of garbage. And extracting information reliably from the web just by using counts popularity is extremely hard. And the only way you can fix this is by intervening. And well, there's two ways to go about it. Right now, they use a bunch of rules. And this is the best, by the way.
[34:44] This is the cream, what you're looking at here, as to what we can do today. This is the sort of thing that Google will use. and so on. likely and a lot of companies use it. Orin Asuncion is a great entrepreneur and he's sold many companies using this type of technology. If we are to go beyond this, there are two routes. One is to start parameterizing these rules. Because right now, linguists tend to create all these rules of syntax and so on.
[35:19] It has to be a noun phrase followed by a verb phrase followed by a noun phrase and so on. Or then they start adding a bunch of exceptions. If you could just parameterize all these and then having a learning machine estimate what are the best, predict labels, and then have people vote yay/nay on labels. If you could create a cool app that allows people to do that, we would learn to do information extraction. a lot better. But if you want to uncover truth, ultimately we need to contend with the concept of causality. We need to actually know what is... we need to conduct an experiment.
[35:56] or we need to actually probe, we need to conduct an action to be able to actually figure out. We would have to do something now, after this point. We would have to click, we would have to do some research, and we would have to decide how we're going to do that research in order to be able to discover truth. So action is necessary for truth. That brings me to the next part.
[36:38] which is-- Learning is not something that just happens where you have a lot of data and then you learn from it. That's how I did it in 340, for those there. But in 540, we're also going to look at when we actually actively go and seek the data, when we do interventions. Because most systems are not just about... being flooded with data and doing something as a result, but you actually do interventions in the world. And most of the data you know, in order to get data we conduct experiments. I mean, this is pretty much at the heart of the scientific method.
[37:20] If we want to understand something we need to perform actions. So, I'm going to do a bit of causality this year in this course because I think it's important. eventually to have true understanding. Okay. The course will also have a lot of neural networks. And recently, there's been a lot of cool stuff happening in an area called deep learning. How many of you saw the Google Frontline deep learning network in the New York Times?
[37:59] One, so machine learning was big in the news. Big enough that two people, three people, four, five managed to see it. It was in the front page of the New York Times. So Google was claiming that they had built this neural network that can look at YouTube videos and just by doing that it can learn what a cat is or what a face is and so on. And so we're going to go carefully over that experiment in this course.
[38:29] state of the art in deep learning. But needless to say, there's been some really amazing advances over the last three years in this area that are bringing us much closer to doing active general object recognition. The speech recognition that got a lot better over the last--at Microsoft and Google, the reason why it got better was because of this deep learning. And so quite a bit of the course, and it's all very biologically inspired. And so we are going to cover that in the course.
[39:03] It's going to be almost like a third of the course. I'm not going to go into neuroscience, but I will not be on this little caricature here on the screen, which is there's a brain. If you're looking at vision information, there's a visual signal that comes into your eye. Again, that visual signal is very tiny, and it's dynamic, and it's about the size of your, the high-resolution bit is the size of your thumb.
[39:37] - At this distance. and then the less, the lower resolution about the size of your hand. But you really don't see much. This is really mostly imagined. Then there is--everything goes by-- this optic nerve to the back of your brain which is the first part of vision and try doing this, try hitting the back of your head. Things go black. There's a reason. That's where your vision is.
[40:10] Over here with LGNAs there's about a million fibers. So that starts giving a rough idea of what the resolution of that image is, thousand by a thousand. And then at the back there, there's an area called V1, which is very interesting. Now I'm going to have to say a few things about that. Now the brain is a good memory device and it's a predictive device. The way we store things in the brain is in an associative manner.
[40:45] And what I mean by that is, for example, something like that, where there's--if you want to retrieve something, when you present a new query, like in this case, a plane with clouds, you retrieve a plane. So it understands that there's a plane there. um Who doesn't mind giving their phone away? Can you say your phone number loud forward? 604-442-7185.
[41:19] Can you say it backwards? I heard 280. - Okay, so significantly slower when he says it backward. And the reason is because you don't really process numbers like a computer does. You don't store numbers, like the, you know, with a random access memory or whatever. What you've learned is this pattern that you repeat, which is 7782391190. So easy, 'cause it's like this, you store that pattern, which is your phone number. You haven't stored the numbers.
[41:55] stored a pattern that sequence same is true with the alphabet a lot of people can have no problem saying the alphabet forward but people struggle saying the alphabet backward because the what they've learned stored is patterns okay sequences and so on so i have a different experience so i'm not a native english speaker but when i'm telling somebody my phone number using chinese pretty fast. So when I'm speaking it using English it takes some time.
[42:33] So my feeling is when I'm speaking my phone number using English, I... the picture of my phone number and read that picture. And I think that might be the part of where time is . Yeah, it's true. But you're doing it by association. So association, by the way, here I did association with time, but it could be spatial. In fact, a lot of people who are very good at memorizing things, it's because they think of an image.
[43:07] Like if I give you a list of 100 objects, then what they try to do, 100 is actually an overkill, 10 objects. So most people would find it really hard to remember 10 objects. But if I, the ones that are good at it, they compose an image with the 10 objects. and then they just recite that image. But they create a pattern and a context, and that makes it easier. somewhat similar to what you're saying. The key point being that memory is associated.
[43:41] one game for today. If you were to pick one of these to be a spider Which one would it be? A jumping spider. How many of you have seen a jumping spider? They're all over the place. So how many people would vote for this first guy being a jumping spider? You only get to pick one, so there's two votes for this one. This one. Zero. Next.
[44:12] Zero. Next. Three volts, four. Next. Four. Next. One, three. This guy. One vote, two, three. Next guy. Okay, so that's like what about 30 Okay. Now...
[44:45] If we compare these numbers, so those were images that were shown to a male jumping spider to see if that male jumping spider would see this as a female that wants to have spider sex. You never expected I was going to pull this one on you. So-- and The point here that I want to illustrate is that even little animals are not that far from us in terms of their ability to recognize us.
[45:15] So spiders are actually able to recognize simple caricatures, drawings. So not just actual images of-- that's a jumping spider. And here you see those jumping spider recognizing can recognize mating behaviors and so on, even when it just of caricatures and so on. So spiders Also, these jumping spiders, they're able to recognize 3D depth. and so on, which is essential to travel in the world. So, we don't need to get to the high-level brains.
[45:48] I started with the high level brains, they're showing a human brain. but the very low level brain of the spider is an extremely sophisticated piece of machinery that We still are struggling to replicate with our own electronic devices. And evolution is beautiful, just like spiders have learned all these amazing-- uh... mental traits, recognition and so on. There's creatures that have evolved, like this moth, that they've evolved to replicate to look like jumping spiders so that they Don't fall prey to them.
[46:26] Um... Most of what we see in the world gets encoded in our brain, but even more, the location of things in the world. Like in this case, what we see in red. So that's a box. And that's square. And there's a raft that moves in the box. And then what we're monitoring is whether a neuron fires or not. And as you can see, there are particular portions periodic interval in the box. where a neuron fires. And that's how you can tell where you are. close your eyes just let that neuron flow.
[46:58] fire tells me where you are. Where you are is encoded in your brain. Now, going back to vision, back to this, All these fibers get into these first level of cells. and V1. And people did the same sort of experiments. there probes. on cats. Thank you. at the back of the brain And then they shortcuts these bars as shown in the screen.
[47:30] vertical bars and horizontal bars. And when you do that, Um... The experiment is by Hubel and Wiesel. You can actually see the whole video in YouTube, if you Google Hubel and Wiesel. Um-- What they found was the following. If-- If the bar is vertical-- is horizontal, sorry-- as shown here in the first drawing. If the bar is horizontal-- that the neuron doesn't fire. The neuron only fires when the bar is at a particular angle.
[48:05] Angle is where the neuron starts firing. That means that each of these neurons is basically tuned to-- vertical, specific lines. Lines are different. orientation. So all these neurons basically have an age map. of the world. because they fire-- different neurons will fire for different edges. Now there's many neurons. there in V1. Each of the little boxes in this collage is one of those neurons. And each of those will-- all 20 million or so.
[48:37] will be monitoring checking whether there's a patch here, whether there's coordinate there. whether there's an edge here. And it's all those features Not just the eyebrows. But it's these 20 million features, when you take compositions of them, that allow us to recognize whether something's a face or not a face. They're all very basic. And it's-- But in having these huge compositions, we're able to do all sorts of recognition. Thank you.
[49:08] So the basic idea-- and we will go over this in the course I'm just... At this stage, I'm just giving an overview which hopefully will inspire you. to continue. Um. And later in the course, we'll actually look at the details of how we build these. uh... But the basic idea is that you will have neurons And each neuron is tuned to one particular stimuli. Like in this case, an edge. that happens in this part of the image. Thank you. So our nature is basically a contrast between light and dark.
[49:41] And when that neuron sees this, it will fire. FIKE gets produced. several spikes start getting produced. If a neuron doesn't see this, it doesn't fight. And this is the basis for constructing a neural network. something that looks at an image and then when you have 2000 neurons connected to the entire image and some of this units will fire, some of them will not fire. The implication of this is that we're mapping an image to a code.
[50:17] of about 2,000 years. We will see in this course how to learn these. how to learn, what are the-- receptive fields for each neuron. what's What should each neuron be paying attention to? What ages? So this we will learn. for each of these neurons We will learn them automatically from data. And then what researchers have done and deep learning is that they then take the output of these neurons and they feed them to a second network.
[50:50] where they learn again the parameters of that network. And they find that they learn things like the subject parts, and so on, when they use images. when they just train it with, say, images of faces. And then further down up here, you have actually entire faces. So, all this gets done automatically these days and this is essentially what Google was Google was taking one of these large networks and was able to just train it And-- And the way it's trained here will be without supervision.
[51:27] So that's going to be an interesting We're not going to be saying these are images of faces are not images of faces. We will learn how to take a huge collection of images And just by using some concepts, which is basically what you imagine should look like what you see. that very basic idea. If I have a learning system that's capable of imagining Um-- what it sees. it should be able to learn interesting concepts. And that's essentially what Google demonstrated. They took a system where you took many images, The idea had been demonstrated to be fair by many groups in academia.
[52:06] But Google just took it to the next level, very large scale. And so they-- by just training on all sorts of images in YouTube neurons that were capable of recognizing faces emerged. So there's one particular neuron in this ensemble of neurons Every time it sees one of these faces, it fires. So now that's also biologically plausible. There's-- paper. There's a very famous story of the Halle Berry Neuron where neuroscience actually conducted a study where you know, Which animal did I use?
[52:42] for that from the chat. Probably a cat. uh... Humans are not allowed in these things. Um-- that that fires-- has a neuron that fires every time this cat sees Halle Berry, a picture of Halle Berry. Thank you. Thank you. and uh... Thank you. And then they can also, in practice, process. switch this neuron on and because you've learned the receptive fields, you can actually learn, "What is the input that makes this neuron fire?"
[53:18] And that was the picture that made it to the front of the New York Times. century. We're going to see how to compute this picture in the rest of the course. I will also see as to how The same idea that you can reconstruct the input can be used in order to do this sort of-- tracking and be able to complete this seen. in order to make this open. protections. Um... Finally, to end this, this is... the speech recognition that works now these days on your phone It does a first level which is it goes from sounds come to central coefficients. That's sort of very easy.
[53:59] - Two, two? Standard. It then uses a deep learning, again, a neural network to go from those basically sound features. to features that can be fed into something called an HMM. in order to do recognition. So throughout the course, we'll go and look at the elements Um-- that are necessary sorry to publish And that's pretty much it. So the next class will start already with--with the--with the covering supervised learning.
[54:30] And we'll start with linear models because they happen to be very easy to understand and I'll let him hold it. Thank you. Bye. Thank you. Thank you. Next time, I want Russia.
[55:10] Yes, yes. The station is a large, multi-use station with a capacity of about 1,000 people. The station is a large, multi-use station with a capacity of about 1,000 people.
[56:10] Thank you. Yeah, I know.
[57:10] Thank you. *Captions by Project readOn* Thank you.
[58:40] I'm just sitting there. I don't think we're in a federal situation.
[59:40] .
Open in the Vidleaf workbench
Search the transcript, select lines, copy quotes with timestamps, translate.
Attribution
"Machine learning - introduction" by Nando de Freitas (https://www.youtube.com/@profnandodf), licensed under CC BY 3.0 (https://creativecommons.org/licenses/by/3.0/). Source video: https://www.youtube.com/watch?v=w2OtwL5T1ow. This page is a text transcript of the video with paragraph breaks and timestamps added; the creator is not affiliated with and does not endorse Vidleaf.
Are you the creator or a rights holder? Request a correction or removal: copyright@vidleaf.app (see About these pages).
Last updated