Artificial Intelligence Applications in Engineering (MUH-920), week 10 of 14

Neural Networks: From the Perceptron to Deep Learning

Prof. Dr. Utku Kose, Süleyman Demirel University

From neurons to perceptrons

McCulloch and Pitts described an idealised neuron that sums weighted binary inputs and fires when the sum exceeds a threshold, and showed that networks of such units can compute logical functions [1]. Rosenblatt's perceptron added learning: After each example, the weights move in the direction that corrects a wrong output, w = w + eta (y - y_hat) x [2]. If the two classes can be separated by a straight line, or a hyperplane in more dimensions, the rule finds a separating line in a finite number of steps.

Minsky and Papert analysed what single-layer perceptrons cannot do [3]. The exclusive or, XOR, is the classic example: The points (0, 0) and (1, 1) belong to one class and (0, 1) and (1, 0) to the other, and no straight line separates them. Engineering data are full of such interactions, for example when a fault appears only if temperature is high and speed is low. The remedy is a hidden layer that transforms the inputs into a space where the classes become separable.

Check your understanding. Why can a single perceptron not learn XOR?

Multilayer networks and backpropagation

A multilayer perceptron stacks layers of units. Each unit computes a weighted sum of the outputs of the previous layer and applies a nonlinear activation function: the logistic sigmoid, the hyperbolic tangent, or the rectified linear unit, ReLU, which returns max(0, z) and eases the training of deep networks [4]. Without nonlinear activations, any stack of layers collapses into one linear map. With them, a network with one sufficiently wide hidden layer can approximate any continuous function on a bounded domain to any accuracy [5, 6], although the theorem says nothing about how many units are needed or whether training finds them.

Training minimises a loss, such as the mean squared error for regression or the cross-entropy for classification, by gradient descent. Backpropagation computes the gradient of the loss with respect to every weight efficiently by applying the chain rule layer by layer from the output back to the input [7]. Each weight then moves a small step against its gradient. Modern frameworks such as PyTorch record the operations of the forward pass and compute these gradients automatically [8].

Animation: A small network learns a curved boundary

A network with one hidden layer is trained by gradient descent on points from two classes. Change the number of hidden units and the learning rate, and watch the boundary bend while the loss falls.

Check your understanding. What does backpropagation compute?

Optimisation in practice

Plain gradient descent uses the whole training set for every step. Stochastic gradient descent uses small random mini-batches, which is faster and adds noise that can help to escape poor regions. Momentum accumulates a moving average of gradients and damps oscillation across narrow valleys. Adam adapts the step size of every weight from running estimates of the first and second moments of its gradient and is a robust default [9]. The learning rate remains the most important setting: Too small and training crawls, too large and the loss oscillates or diverges.

Initialisation matters because signals must neither vanish nor explode as they pass through many layers. Glorot and Bengio proposed scaling the initial weights with the number of inputs and outputs of a layer [10]. Batch normalisation standardises the inputs of a layer over each mini-batch and often allows larger learning rates [11]. Inputs of different units, such as temperatures in degrees and pressures in millibars, should be standardised before training, as for k-means and support vector machines.

Generalisation and regularisation

Networks with many weights can memorise their training data. The standard defences are a validation set with early stopping, which keeps the weights from the epoch with the lowest validation loss; weight decay, which penalises large weights like ridge regression; and dropout, which randomly switches off units during training so that the network cannot rely on any single unit [12]. Learning curves of training and validation loss over epochs show whether a network underfits, with both losses high, or overfits, with a growing gap.

On tabular data of moderate size, the tree ensembles of Weeks 7 and 8 are often as accurate as networks and easier to tune, so a network must earn its place by a fair comparison. Networks have long been applied in engineering: Yeh modelled concrete strength with them [13], and metaheuristics such as the ant lion optimiser have been used to train them where gradient methods struggle [14]. Their decisive advantages appear with images, signals, sequences and text, where they learn features that would otherwise have to be designed by hand [15]. Weeks 11 to 13 follow this path.

Check your understanding. Training loss keeps falling, while validation loss has been rising for 30 epochs. What should be done?

Python step 10: Matrices, shapes and tensors

The Python step of this week is part of the Colab notebook, where every explanation stands next to a cell that runs it and the step closes with a quick check and exercises with immediate feedback. The printable lecture notes contain the same step together with the outputs of its code.

Open Python step 10 in Colab View the notebook on GitHub

Review cards

Select a card to turn it over.

Perceptron
A threshold unit with learned weights; converges on linearly separable data [2].
Backpropagation
Chain-rule computation of all weight gradients in one backward pass [7].
ReLU
max(0, z); a simple activation that eases training of deep networks [4].
Adam
Adaptive moment estimation; per-weight step sizes from running gradient moments [9].
Early stopping
Keep the weights from the epoch with the lowest validation loss.
nn.Module
PyTorch base class for networks: layers in __init__, computation in forward.

Continue the week

The week continues with the simulation and the self-assessment of the interactive lab and with the Python step and the hands-on work of the Colab notebook. The week overview lists the discipline challenges, the weekly task and the research assignment.

Interactive lab Colab notebook Self-assessment Week overview and tasks

References

[1] McCulloch, W. S., & Pitts, W. (1943). A logical calculus of the ideas immanent in nervous activity. The Bulletin of Mathematical Biophysics, 5(4), 115-133. https://doi.org/10.1007/BF02478259

[2] Rosenblatt, F. (1958). The perceptron: A probabilistic model for information storage and organization in the brain. Psychological Review, 65(6), 386-408. https://doi.org/10.1037/h0042519

[3] Minsky, M., & Papert, S. (1969). Perceptrons: An Introduction to Computational Geometry. MIT Press.

[4] Nair, V., & Hinton, G. E. (2010). Rectified linear units improve restricted Boltzmann machines. In Proceedings of the 27th International Conference on Machine Learning (pp. 807-814).

[5] Cybenko, G. (1989). Approximation by superpositions of a sigmoidal function. Mathematics of Control, Signals and Systems, 2(4), 303-314. https://doi.org/10.1007/BF02551274

[6] Hornik, K., Stinchcombe, M., & White, H. (1989). Multilayer feedforward networks are universal approximators. Neural Networks, 2(5), 359-366. https://doi.org/10.1016/0893-6080(89)90020-8

[7] Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533-536. https://doi.org/10.1038/323533a0

[8] Paszke, A., Gross, S., Massa, F., et al. (2019). PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32 (pp. 8024-8035). https://arxiv.org/abs/1912.01703

[9] Kingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations (ICLR 2015). https://arxiv.org/abs/1412.6980

[10] Glorot, X., & Bengio, Y. (2010). Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, PMLR 9 (pp. 249-256).

[11] Ioffe, S., & Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the 32nd International Conference on Machine Learning, PMLR 37 (pp. 448-456). https://arxiv.org/abs/1502.03167

[12] Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. Journal of Machine Learning Research, 15(56), 1929-1958.

[13] Yeh, I.-C. (1998). Modeling of strength of high-performance concrete using artificial neural networks. Cement and Concrete Research, 28(12), 1797-1808. https://doi.org/10.1016/S0008-8846(98)00165-3

[14] Kose, U. (2018). An ant-lion optimizer-trained artificial neural network system for chaotic electroencephalogram (EEG) prediction. Applied Sciences, 8(9), 1613. https://doi.org/10.3390/app8091613

[15] LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436-444. https://doi.org/10.1038/nature14539