News classification
Contact us
- Add: No. 9, North Fourth Ring Road, Haidian District, Beijing. It mainly includes face recognition, living detection, ID card recognition, bank card recognition, business card recognition, license plate recognition, OCR recognition, and intelligent recognition technology.
- Tel: 13146317170 廖经理
- Fax:
- Email: 398017534@qq.com
The history of artificial intelligence AI
The history of artificial intelligence AI
Artificial intelligence is the application of intelligent learning algorithms that use the experience of a large amount of data to improve the performance of the system itself, thus accomplishing what human intelligence can do. It can be said that artificial intelligence is inseparable from data analysis and machine learning. The theory and methods of analyzing intelligent data analysis have become one of the necessary foundations of artificial intelligence. Artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and consumes a new kind of intelligent machine that can respond in a similar way to human intelligence. The discussion of this category includes robotics, speech recognition, image recognition, Natural speech disposal and expert systems. Since the birth of artificial intelligence, the theory and technology have become more and more mature, and the application scope has expanded from time to time. It can be imagined that the technological products brought by artificial intelligence in the future will be the "containers" of human intelligence. Artificial intelligence can stop imitating people's understanding and thought process. Artificial intelligence is not human intelligence, but it can be considered like human beings, and it may transcend human intelligence.
The beginning of artificial intelligence has been unable to circumvent Alan Mathison Turing (June 23, 1912 - June 7, 1954) as a hero of World War II, a huge mathematician. , cryptographer, logician, he has common ideas, general thoughts, legendary experience (his personal experience can see the 2014 movie "simulation game", a hero behind the success of war).
In 1950, Alan Messon Turing raised questions about machine thinking. His paper "Computingmachiery and intelligence" caused widespread attention and far-reaching influence. In October 1950, Turing published a paper. "Can the machine be considered?" This period of time makes Turing win the title of "father of artificial intelligence." It also announced the arrival of the artificial intelligence (AI) period.
Of course, the development of things is by no means an occasional. The arrival of the AI period must have historical inevitability. We will not talk about the very old history. Starting from Turing, we will introduce the evolution of technology through the trajectory of time. In the following articles, we will introduce some of the key people in detail and let us worship the gods.
1. The “inference period” of artificial intelligence in the 1950s and 1970s.
Hepburn opened the first step in machine learning in 1949 based on a neuropsychological learning mechanism. It is later called the Hebb learning rule. The Hebb learning rule is a no-monitoring learning rule. The result of this learning is that the network can extract the statistical characteristics of the exercise set, thereby dividing the input information into several categories according to their similarity levels. This is in line with the process of human beings seeing and understanding the world. Human beings see and recognize the world at a considerable level and stop classifying according to the statistical characteristics of things.
As can be seen from the above formula, the amount of weight adjustment is proportional to the product of input and output. Obviously, the form often presented will have a greater impact on the weight vector. Under this circumstance, the Hebb learning rule needs to pre-set the saturation value to avoid the unweighted growth of the weight when the input and output are positive and negative.
The Hebb learning rules differ from the "conditioned reflex" mechanism and have been proven by neuronal theory. For example, Pavlov's conditioning experiment: every time the dog is fed, it will ring before feeding. After a long time, the dog will connect the ringtone with the food. In the future, if you ring but don't give food, the dog will drool (of course this experiment is also thought to be the beginning of advertising, neurology).
In 1950, Alan Turing invented the Turing test to determine whether the computer could be intelligent. The Turing test assumes that if a machine can talk to humans (via telex equipment) and cannot be distinguished from its machine identity, then the machine is said to be intelligent. This simplification makes it possible for Turing to convincingly clarify that "the machine under consideration" is possible.
On June 8, 2014, a chat robot victory called Eugene Gustman made humans believe that it was a 13-year-old boy and became the first Turing-tested computer ever. This is thought to be a milestone in artificial intelligence. But this is 14 years later than Turing himself predicted.
In the summer of 1956, a group of forward-looking young scientists, led by McCarthy, Minsky, Rochester and Shennong, gathered together to discuss and discuss a series of related problems in the use of machines to imitate intelligence, and proposed for the first time. The term "artificial intelligence", which marks the official birth of the emerging discipline of "artificial intelligence."
In 1957, Rosen Blatter proposed a second model based on the neuro-perceptual science background, which is very similar to today's machine learning model. This was a very exciting discovery at the time, and it was more applicable than Herb's idea. Based on this model, Rosen Blatter designed the first computer neural network, the perceptron, which mimics how the human brain works. Rosen Blatter defines the perceptron as follows:
The perceptron is designed to clarify some of the fundamental attributes of a general intelligent system. It is not constrained by individual practices or things that are not normally understood, nor is it disrupted by the condition of individual biological organisms.
The perceptron is designed to illustrate some of the fundamental properties of intelligent systems in general, without becoming too deeply enmeshed in the special, and frequently unknown, conditions which hold for particular biological organisms.
Three years later, Vidro first used the Delta Learning Rule (ie, Least Squares) for the perceptron's exercise steps and invented a good linear classifier.
In 1967, the nearest neighbor algorithm was presented to enable the computer to stop simple form recognition. The central idea of the kNN algorithm is that if the majority of the k most neighboring samples in a feature space belong to a certain category, the sample also belongs to this category and has the characteristics of the samples on this category. This is the so-called "minority to follow the majority" standard.
In 1969, Marvin Minsky proposed the famous XOR problem, pointing out that the perceptron is ineffective in linearly inseparable data distribution. The researchers of the neural network entered the cold winter and did not recover until 1980.
From the mid-1960s to the end of the 1970s, the pace of machine learning was stagnation. Whether it is theoretical research or computer hardware limitations, the entire artificial intelligence category has encountered a lot of bottlenecks. Although the structure-based learning system of Winston and the logic-based learning system of Hayes Roth in this period have been largely paused, they can only learn a single concept and fail to put into practical application. .
During this period, it was generally thought that only the machine was given logical reasoning to complete artificial intelligence. However, people later discovered that only with logical reasoning, the machine is far from being intelligent. In the late 1960s and early 1970s, the neural network learning machine turned to a low tide due to the theoretical defects that failed to reach the expected effect. The purpose of this period of research is to imitate the concept learning process of human beings, and to use logical construction or graph construction as the internal depiction of the machine. Machines can use symbols to describe concepts (symbolic concept acquisition) and propose various assumptions about learning concepts. In fact, the entire AI category encountered bottlenecks during this period. The limited memory and processing speed of the computer at the time was lacking to handle any practical AI issues. The requester had a child-level understanding of the world, and the researchers quickly discovered that the request was too high: no one could make such a grand database in 1970, and no one knew how a program could learn such a rich message. It can be said that the lag of computing talent has led to the stagnation of AI technology. Although academic thoughts have reached a high level at this stage, the development of industry is seriously lagging behind. Science and technology complement each other, and the delay of either side will drag the other side. Carry out.
2. The period of the "study project" of artificial intelligence in the 1970s and 1990s.
During this period, people thought that to make the machine smart, they should try to let the machine learn to learn, so the computer system got a lot of development. Later, people found that it was quite difficult to summarize the knowledge and instill it in the computer. For example, if you want to develop an artificial intelligence system for disease diagnosis, first find a doctor with experience to summarize the laws and knowledge of the disease, and then let the machine stop learning, but at the stage of the summary of the study, it has cost a lot. Labor costs, the machine is just an automated tool to execute the library of knowledge, can not reach the true degree of intelligence and replace human work.
In the 1981 neural network back propagation (BP) algorithm, Weibos proposed a multi-layer perceptron model (known as "artificial neural network", ANN). Although the BP algorithm was proposed in the name of "reverse mode of automatic differentiation" as early as 1970, it did not really work until then, and until now BP algorithm is still a neural network architecture. The key element. With these new ideas, the discussion of neural networks has accelerated.
In 1985-1986, the neural network researchers (Rummel Hart, Hinton, Williams-Her, Nielsen) successively proposed the concept of multi-parameter linear programming (MLP) using BP algorithm to become deep learning later. Cornerstone.
In another pedigree, Quinlan proposed a well-known machine learning algorithm in 1986, which we call the “decision tree”, in more detail the ID3 algorithm. This is the breaking point of another mainstream machine learning algorithm. In addition, the ID3 algorithm has been released as a software that can find more ideal cases with simple planning and clear inferences, which is exactly the opposite of the neural network black box model.
A decision tree is a predictive model that represents a mapping between object properties and object values. Each node in the tree represents an object, and each forked path represents a possible attribute value, and each leaf node corresponds to the object represented by the path from the root node to the leaf node. value. The decision tree has only a single output. If you want to have a complex output, you can set up an independent decision tree to handle different outputs. Decision tree in data mining is a technique that is often used to analyze data and can also be used for forecasting.
After the ID3 algorithm was proposed, the discussion community has explored many different choices or improvements (such as ID4, regression tree, CART algorithm, etc.), which are still active in the field of machine learning.
In 1990, Schapire first constructed a polynomial-level algorithm, which was the original Boosting algorithm. A year later, Freund proposed a more efficient Boosting algorithm. However, there are common theoretical shortcomings between the two algorithms, that is, they all request that the weak learning algorithm learns the correct lower limit beforehand.
In 1992, the representation of support vector machine (SVM) was another important breakthrough in the field of machine learning. The algorithm has a very powerful theoretical position and empirical results. In the 1990s, the whole data mining, machine learning, and form identification were popular. It is still very hot. At that time, the machine learning seminar was divided into two groups: NN and SVM. However, after the support vector machine with kernel function was proposed around 2000. SVM has achieved better results in many of the tasks previously occupied by NN. In addition, SVM related to NN can also apply all the deep knowledge about convex optimization, generalized edge theory and kernel function. Therefore, SVM can promote theoretical and theoretical improvements from different disciplines (of course, SVM still has strong vitality, and SVM is still a good tool when deep learning algorithms cannot handle problems).
Neural networks and support vector machines are constantly in a "competitive" relationship. SVM applies the expansion theorem of kernel function, no need to know the explicit expression of nonlinear mapping; because it is a linear learning machine in high-dimensional feature space, compared with linear model, it not only increases the computational complexity, but also At some level, the "dimensional disaster" is prevented (about the meaning that the quality of the high-dimensional space will be infinitely close to his appearance, what a fearful thing, your skin is broken and you will find that you are empty).
The neural network has been questioned again. After the research by Hochreiter et al. in 1991 and Hochreiter et al. in 2001, it is indicated that when the BP algorithm is applied, the NN neurons will show a gradient loss after saturation. Simply put, the earlier neural network algorithm is easier to exercise, a lot of experience parameter setting; the exercise speed is slower, and the effect is not better than other methods when the level is less (less than or equal to 3), so this period NN is superior to SVM.
Third, from 2000 to the present, the period of "data mining" of artificial intelligence.
With the introduction and application of various machine learning algorithms, especially the deep learning technology, people hope that the machine can analyze a large amount of data to automatically learn the knowledge and complete the intelligence. During this period, with the improvement of computer hardware and the development of big data profiling technology, the degree of machine acquisition, storage, and disposal of data has greatly improved. In particular, the deep learning technology has a great understanding of learning than the previous shallow learning. Alpha Go and the masters of Chinese and Korean Go chess have been one of the high-level representatives of artificial intelligence.
In 2001, the decision tree model was proposed by Dr. Brayman. It is an algorithm that integrates multiple trees through the idea of integrated learning. Its fundamental unit is the decision tree, and its essence belongs to a branch of machine learning. Ensemble Learning approach. There are two keywords in the title of random forest, one is “random” and the other is “forest”. "Forest" We know very well that a tree is called a tree, so hundreds of thousands of trees can be called forests. This kind of metaphor is very appropriate. In fact, this is also the main idea of the random forest - the expression of integrated thinking.
In fact, from an intuitive point of view, each decision tree is a classifier (assuming that it is now a classification problem), then for an input sample, N trees will have N classification results. The random forest integrates all the classified voting results, and specifies the category with the most votes as the final output. This is one.
The beginning of artificial intelligence has been unable to circumvent Alan Mathison Turing (June 23, 1912 - June 7, 1954) as a hero of World War II, a huge mathematician. , cryptographer, logician, he has common ideas, general thoughts, legendary experience (his personal experience can see the 2014 movie "simulation game", a hero behind the success of war).
In 1950, Alan Messon Turing raised questions about machine thinking. His paper "Computingmachiery and intelligence" caused widespread attention and far-reaching influence. In October 1950, Turing published a paper. "Can the machine be considered?" This period of time makes Turing win the title of "father of artificial intelligence." It also announced the arrival of the artificial intelligence (AI) period.
Of course, the development of things is by no means an occasional. The arrival of the AI period must have historical inevitability. We will not talk about the very old history. Starting from Turing, we will introduce the evolution of technology through the trajectory of time. In the following articles, we will introduce some of the key people in detail and let us worship the gods.
1. The “inference period” of artificial intelligence in the 1950s and 1970s.
Hepburn opened the first step in machine learning in 1949 based on a neuropsychological learning mechanism. It is later called the Hebb learning rule. The Hebb learning rule is a no-monitoring learning rule. The result of this learning is that the network can extract the statistical characteristics of the exercise set, thereby dividing the input information into several categories according to their similarity levels. This is in line with the process of human beings seeing and understanding the world. Human beings see and recognize the world at a considerable level and stop classifying according to the statistical characteristics of things.
As can be seen from the above formula, the amount of weight adjustment is proportional to the product of input and output. Obviously, the form often presented will have a greater impact on the weight vector. Under this circumstance, the Hebb learning rule needs to pre-set the saturation value to avoid the unweighted growth of the weight when the input and output are positive and negative.
The Hebb learning rules differ from the "conditioned reflex" mechanism and have been proven by neuronal theory. For example, Pavlov's conditioning experiment: every time the dog is fed, it will ring before feeding. After a long time, the dog will connect the ringtone with the food. In the future, if you ring but don't give food, the dog will drool (of course this experiment is also thought to be the beginning of advertising, neurology).
In 1950, Alan Turing invented the Turing test to determine whether the computer could be intelligent. The Turing test assumes that if a machine can talk to humans (via telex equipment) and cannot be distinguished from its machine identity, then the machine is said to be intelligent. This simplification makes it possible for Turing to convincingly clarify that "the machine under consideration" is possible.
On June 8, 2014, a chat robot victory called Eugene Gustman made humans believe that it was a 13-year-old boy and became the first Turing-tested computer ever. This is thought to be a milestone in artificial intelligence. But this is 14 years later than Turing himself predicted.
In the summer of 1956, a group of forward-looking young scientists, led by McCarthy, Minsky, Rochester and Shennong, gathered together to discuss and discuss a series of related problems in the use of machines to imitate intelligence, and proposed for the first time. The term "artificial intelligence", which marks the official birth of the emerging discipline of "artificial intelligence."
In 1957, Rosen Blatter proposed a second model based on the neuro-perceptual science background, which is very similar to today's machine learning model. This was a very exciting discovery at the time, and it was more applicable than Herb's idea. Based on this model, Rosen Blatter designed the first computer neural network, the perceptron, which mimics how the human brain works. Rosen Blatter defines the perceptron as follows:
The perceptron is designed to clarify some of the fundamental attributes of a general intelligent system. It is not constrained by individual practices or things that are not normally understood, nor is it disrupted by the condition of individual biological organisms.
The perceptron is designed to illustrate some of the fundamental properties of intelligent systems in general, without becoming too deeply enmeshed in the special, and frequently unknown, conditions which hold for particular biological organisms.
Three years later, Vidro first used the Delta Learning Rule (ie, Least Squares) for the perceptron's exercise steps and invented a good linear classifier.
In 1967, the nearest neighbor algorithm was presented to enable the computer to stop simple form recognition. The central idea of the kNN algorithm is that if the majority of the k most neighboring samples in a feature space belong to a certain category, the sample also belongs to this category and has the characteristics of the samples on this category. This is the so-called "minority to follow the majority" standard.
In 1969, Marvin Minsky proposed the famous XOR problem, pointing out that the perceptron is ineffective in linearly inseparable data distribution. The researchers of the neural network entered the cold winter and did not recover until 1980.
From the mid-1960s to the end of the 1970s, the pace of machine learning was stagnation. Whether it is theoretical research or computer hardware limitations, the entire artificial intelligence category has encountered a lot of bottlenecks. Although the structure-based learning system of Winston and the logic-based learning system of Hayes Roth in this period have been largely paused, they can only learn a single concept and fail to put into practical application. .
During this period, it was generally thought that only the machine was given logical reasoning to complete artificial intelligence. However, people later discovered that only with logical reasoning, the machine is far from being intelligent. In the late 1960s and early 1970s, the neural network learning machine turned to a low tide due to the theoretical defects that failed to reach the expected effect. The purpose of this period of research is to imitate the concept learning process of human beings, and to use logical construction or graph construction as the internal depiction of the machine. Machines can use symbols to describe concepts (symbolic concept acquisition) and propose various assumptions about learning concepts. In fact, the entire AI category encountered bottlenecks during this period. The limited memory and processing speed of the computer at the time was lacking to handle any practical AI issues. The requester had a child-level understanding of the world, and the researchers quickly discovered that the request was too high: no one could make such a grand database in 1970, and no one knew how a program could learn such a rich message. It can be said that the lag of computing talent has led to the stagnation of AI technology. Although academic thoughts have reached a high level at this stage, the development of industry is seriously lagging behind. Science and technology complement each other, and the delay of either side will drag the other side. Carry out.
2. The period of the "study project" of artificial intelligence in the 1970s and 1990s.
During this period, people thought that to make the machine smart, they should try to let the machine learn to learn, so the computer system got a lot of development. Later, people found that it was quite difficult to summarize the knowledge and instill it in the computer. For example, if you want to develop an artificial intelligence system for disease diagnosis, first find a doctor with experience to summarize the laws and knowledge of the disease, and then let the machine stop learning, but at the stage of the summary of the study, it has cost a lot. Labor costs, the machine is just an automated tool to execute the library of knowledge, can not reach the true degree of intelligence and replace human work.
In the 1981 neural network back propagation (BP) algorithm, Weibos proposed a multi-layer perceptron model (known as "artificial neural network", ANN). Although the BP algorithm was proposed in the name of "reverse mode of automatic differentiation" as early as 1970, it did not really work until then, and until now BP algorithm is still a neural network architecture. The key element. With these new ideas, the discussion of neural networks has accelerated.
In 1985-1986, the neural network researchers (Rummel Hart, Hinton, Williams-Her, Nielsen) successively proposed the concept of multi-parameter linear programming (MLP) using BP algorithm to become deep learning later. Cornerstone.
In another pedigree, Quinlan proposed a well-known machine learning algorithm in 1986, which we call the “decision tree”, in more detail the ID3 algorithm. This is the breaking point of another mainstream machine learning algorithm. In addition, the ID3 algorithm has been released as a software that can find more ideal cases with simple planning and clear inferences, which is exactly the opposite of the neural network black box model.
A decision tree is a predictive model that represents a mapping between object properties and object values. Each node in the tree represents an object, and each forked path represents a possible attribute value, and each leaf node corresponds to the object represented by the path from the root node to the leaf node. value. The decision tree has only a single output. If you want to have a complex output, you can set up an independent decision tree to handle different outputs. Decision tree in data mining is a technique that is often used to analyze data and can also be used for forecasting.
After the ID3 algorithm was proposed, the discussion community has explored many different choices or improvements (such as ID4, regression tree, CART algorithm, etc.), which are still active in the field of machine learning.
In 1990, Schapire first constructed a polynomial-level algorithm, which was the original Boosting algorithm. A year later, Freund proposed a more efficient Boosting algorithm. However, there are common theoretical shortcomings between the two algorithms, that is, they all request that the weak learning algorithm learns the correct lower limit beforehand.
In 1992, the representation of support vector machine (SVM) was another important breakthrough in the field of machine learning. The algorithm has a very powerful theoretical position and empirical results. In the 1990s, the whole data mining, machine learning, and form identification were popular. It is still very hot. At that time, the machine learning seminar was divided into two groups: NN and SVM. However, after the support vector machine with kernel function was proposed around 2000. SVM has achieved better results in many of the tasks previously occupied by NN. In addition, SVM related to NN can also apply all the deep knowledge about convex optimization, generalized edge theory and kernel function. Therefore, SVM can promote theoretical and theoretical improvements from different disciplines (of course, SVM still has strong vitality, and SVM is still a good tool when deep learning algorithms cannot handle problems).
Neural networks and support vector machines are constantly in a "competitive" relationship. SVM applies the expansion theorem of kernel function, no need to know the explicit expression of nonlinear mapping; because it is a linear learning machine in high-dimensional feature space, compared with linear model, it not only increases the computational complexity, but also At some level, the "dimensional disaster" is prevented (about the meaning that the quality of the high-dimensional space will be infinitely close to his appearance, what a fearful thing, your skin is broken and you will find that you are empty).
The neural network has been questioned again. After the research by Hochreiter et al. in 1991 and Hochreiter et al. in 2001, it is indicated that when the BP algorithm is applied, the NN neurons will show a gradient loss after saturation. Simply put, the earlier neural network algorithm is easier to exercise, a lot of experience parameter setting; the exercise speed is slower, and the effect is not better than other methods when the level is less (less than or equal to 3), so this period NN is superior to SVM.
Third, from 2000 to the present, the period of "data mining" of artificial intelligence.
With the introduction and application of various machine learning algorithms, especially the deep learning technology, people hope that the machine can analyze a large amount of data to automatically learn the knowledge and complete the intelligence. During this period, with the improvement of computer hardware and the development of big data profiling technology, the degree of machine acquisition, storage, and disposal of data has greatly improved. In particular, the deep learning technology has a great understanding of learning than the previous shallow learning. Alpha Go and the masters of Chinese and Korean Go chess have been one of the high-level representatives of artificial intelligence.
In 2001, the decision tree model was proposed by Dr. Brayman. It is an algorithm that integrates multiple trees through the idea of integrated learning. Its fundamental unit is the decision tree, and its essence belongs to a branch of machine learning. Ensemble Learning approach. There are two keywords in the title of random forest, one is “random” and the other is “forest”. "Forest" We know very well that a tree is called a tree, so hundreds of thousands of trees can be called forests. This kind of metaphor is very appropriate. In fact, this is also the main idea of the random forest - the expression of integrated thinking.
In fact, from an intuitive point of view, each decision tree is a classifier (assuming that it is now a classification problem), then for an input sample, N trees will have N classification results. The random forest integrates all the classified voting results, and specifies the category with the most votes as the final output. This is one.