Learning to act
In Week 1, an engineer wrote the rule of the thermostat. Reinforcement learning lets an agent find the rule itself. At every step the agent observes a state, chooses an action and receives a reward and a new state; its goal is a policy that maximises the expected sum of discounted future rewards [1]. The formal model is a Markov decision process: The next state depends only on the current state and action. The discount factor gamma, between 0 and 1, sets how far the agent looks ahead. Unlike supervised learning, nobody tells the agent the correct action; it must explore to find out which actions pay off, and exploit what it has learned to collect reward.
The value of taking action a in state s and acting well afterwards is Q(s, a). Q-learning improves an estimate of Q after every step with Q(s, a) = Q(s, a) + alpha (r + gamma max Q(s', a') - Q(s, a)), where alpha is a learning rate and the maximum runs over the actions in the next state; Watkins and Dayan proved that it converges to the optimal values under suitable conditions [2]. An epsilon-greedy policy explores with a random action with probability epsilon and exploits the best known action otherwise. Tables work for small discrete problems; for large or continuous state spaces, a neural network approximates Q, as in the deep Q-network that learned to play Atari games from pixels [3].
Animation: Q-learning on a cliff walk
The agent starts at the lower left and must reach the lower right without stepping into the cliff, which costs heavily. Each step costs a little. Watch the arrows of the greedy policy and the colours of the state values change as episodes accumulate. Raise epsilon to explore more.
Check your understanding. Q(s, a) = 2.0, the reward is -1, gamma = 0.9, the best Q-value in the next state is 3.0 and alpha = 0.5. What is the updated Q(s, a)?
Reinforcement learning in engineering
Control problems were among the first testbeds of reinforcement learning: The cart-pole balancing task goes back to adaptive elements studied in the early 1980s [4], and Gymnasium now provides it with a standard interface for experiments [5]. Robotics has used reinforcement learning for manipulation and locomotion [6], and deep reinforcement learning has controlled the magnetic coils of a tokamak in a fusion experiment [7]. These successes share a feature: The agent learned mostly in simulation. Exploration on a real plant can damage equipment or endanger people, the reward may be satisfied in unintended ways, and a policy trained under one set of conditions may fail under others; Amodei and colleagues list safe exploration, reward hacking and distributional shift among the concrete problems of AI safety [8]. Simulators, constraints on actions, human oversight and gradual deployment are therefore part of any engineering use.
Check your understanding. Why do most engineering applications of reinforcement learning train the agent in simulation first?
Physics-informed learning
Engineering rarely starts from zero knowledge. Conservation laws, constitutive equations and boundary conditions are known even when parameters or source terms are not. Physics-informed neural networks add the residual of the governing differential equation, evaluated at many collocation points, to the loss of a network whose input is space or time [9]. Automatic differentiation provides the derivatives of the network output exactly. The network then fits the data where they exist and obeys the physics everywhere else, which makes it far more reliable outside the measured range than a data-only network. The same framework solves inverse problems, estimating unknown coefficients of the equation from data. Karniadakis and colleagues review the broader field of physics-informed machine learning [10]; related methods discover governing equations from data by sparse regression [11, 12]. Physics-informed networks can be hard to train for stiff or multiscale problems, and classical numerical solvers remain the reference where they apply.

Check your understanding. What does the physics term in the loss of a physics-informed network contain?
Digital twins
Grieves and Vickers described a digital twin as a virtual representation of a physical product that is linked to it by data throughout its life, so that behaviour can be predicted and problems anticipated [13]. Tao and colleagues surveyed its industrial use in design, production and maintenance [14]. A twin combines a model of the asset, a data connection that keeps the model's state and parameters current, and services such as monitoring, forecasting, what-if analysis and controller testing. System identification, estimating model parameters from measured inputs and outputs, is its simplest learning component; the notebook identifies the thermal parameters of the room of Week 1 from noisy data and uses the resulting model to forecast and to test a controller. The methods of the whole course appear in twins: rules and fuzzy logic for decisions, optimisation for calibration, machine learning for components without physical models, anomaly detection for monitoring and reinforcement learning for control.
Responsible engineering of AI systems
Engineering codes of ethics require practitioners to hold public safety paramount, and they apply to systems that include learned components. Earlier work on machine ethics and AI safety argued that safety must be designed in rather than added after deployment [15]. Regulation now makes parts of this explicit. The European Union's Artificial Intelligence Act, Regulation (EU) 2024/1689, follows a risk-based approach: A few practices are prohibited, such as social scoring and the inference of emotions in the workplace and in education except for medical or safety reasons; high-risk systems, which include safety components in the management of critical infrastructure such as water, gas, heating and electricity supply, and systems used in employment and education, must meet requirements on risk management, data governance, documentation, human oversight, accuracy and robustness; and systems such as chatbots and generators of synthetic content carry transparency obligations [16].
Frameworks turn principles into practice. The NIST AI Risk Management Framework organises activities into four functions: Govern establishes policies, roles and accountability; Map establishes the context and identifies risks; Measure analyses and tracks risks with appropriate metrics; and Manage prioritises and acts on them [17]. ISO/IEC 42001 specifies a management system for organisations that develop or use AI, in the style of other management system standards [18]. Türkiye's national strategy places trustworthy and responsible AI among its priorities [19]. The energy and resource costs of training and running models are part of the same responsibility [20].
Check your understanding. Under the EU AI Act, which category most likely applies to an AI system that acts as a safety component in the management of a city's water supply network?
The course in retrospect
The fourteen weeks form one toolbox. Search plans routes and sequences; rules encode standards; fuzzy logic turns words into control; evolutionary and swarm algorithms optimise designs; machine learning predicts, classifies and detects anomalies; deep learning perceives images and sequences; language models work with text and code; and reinforcement learning, physics-informed learning and digital twins bring learning into operation. The engineering skill lies in choosing the simplest tool that meets the requirement, validating it honestly, and knowing its limits. The final project, described in the exams folder of the repository, asks for exactly this on a problem of the student's own field.
Python step 14: Environments, learning and organised code
The Python step of this week is part of the Colab notebook, where every explanation stands next to a cell that runs it and the step closes with a quick check and exercises with immediate feedback. The printable lecture notes contain the same step together with the outputs of its code.
Review cards
Select a card to turn it over.
Continue the week
The week continues with the simulation and the self-assessment of the interactive lab and with the Python step and the hands-on work of the Colab notebook. The week overview lists the discipline challenges, the weekly task and the research assignment.
Interactive lab Colab notebook Self-assessment Week overview and tasks
References
[1] Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.
[2] Watkins, C. J. C. H., & Dayan, P. (1992). Q-learning. Machine Learning, 8(3-4), 279-292. https://doi.org/10.1007/BF00992698
[3] Mnih, V., Kavukcuoglu, K., Silver, D., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529-533. https://doi.org/10.1038/nature14236
[4] Barto, A. G., Sutton, R. S., & Anderson, C. W. (1983). Neuronlike adaptive elements that can solve difficult learning control problems. IEEE Transactions on Systems, Man, and Cybernetics, SMC-13(5), 834-846. https://doi.org/10.1109/TSMC.1983.6313077
[5] Towers, M., Kwiatkowski, A., Terry, J., et al. (2024). Gymnasium: A standard interface for reinforcement learning environments. arXiv preprint arXiv:2407.17032. https://arxiv.org/abs/2407.17032
[6] Kober, J., Bagnell, J. A., & Peters, J. (2013). Reinforcement learning in robotics: A survey. The International Journal of Robotics Research, 32(11), 1238-1274. https://doi.org/10.1177/0278364913495721
[7] Degrave, J., Felici, F., Buchli, J., et al. (2022). Magnetic control of tokamak plasmas through deep reinforcement learning. Nature, 602(7897), 414-419. https://doi.org/10.1038/s41586-021-04301-9
[8] Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete problems in AI safety. arXiv preprint arXiv:1606.06565. https://arxiv.org/abs/1606.06565
[9] Raissi, M., Perdikaris, P., & Karniadakis, G. E. (2019). Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378, 686-707. https://doi.org/10.1016/j.jcp.2018.10.045
[10] Karniadakis, G. E., Kevrekidis, I. G., Lu, L., Perdikaris, P., Wang, S., & Yang, L. (2021). Physics-informed machine learning. Nature Reviews Physics, 3(6), 422-440. https://doi.org/10.1038/s42254-021-00314-5
[11] Brunton, S. L., Proctor, J. L., & Kutz, J. N. (2016). Discovering governing equations from data by sparse identification of nonlinear dynamical systems. Proceedings of the National Academy of Sciences, 113(15), 3932-3937. https://doi.org/10.1073/pnas.1517384113
[12] Brunton, S. L., & Kutz, J. N. (2022). Data-Driven Science and Engineering: Machine Learning, Dynamical Systems, and Control (2nd ed.). Cambridge University Press. https://doi.org/10.1017/9781009089517
[13] Grieves, M., & Vickers, J. (2017). Digital twin: Mitigating unpredictable, undesirable emergent behavior in complex systems. In Transdisciplinary Perspectives on Complex Systems (F.-J. Kahlen, S. Flumerfelt, & A. Alves, Eds.) (pp. 85-113). Springer. https://doi.org/10.1007/978-3-319-38756-7_4
[14] Tao, F., Zhang, H., Liu, A., & Nee, A. Y. C. (2019). Digital twin in industry: State-of-the-art. IEEE Transactions on Industrial Informatics, 15(4), 2405-2415. https://doi.org/10.1109/TII.2018.2873186
[15] Kose, U. (2018). Are we safe enough in the future of artificial intelligence? A discussion on machine ethics and artificial intelligence safety. BRAIN. Broad Research in Artificial Intelligence and Neuroscience, 9(2), 184-197.
[16] European Parliament and Council of the European Union (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union, L series, 12 July 2024. https://eur-lex.europa.eu/eli/reg/2024/1689/oj
[17] National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. https://doi.org/10.6028/NIST.AI.100-1
[18] International Organization for Standardization (2023). ISO/IEC 42001:2023 Information technology, Artificial intelligence, Management system. ISO.
[19] Presidency of the Republic of Türkiye Digital Transformation Office, & Ministry of Industry and Technology (2021). National Artificial Intelligence Strategy 2021-2025 (Ulusal Yapay Zekâ Stratejisi 2021-2025). Ankara. Action plan updated for 2024-2025. https://www.cbddo.gov.tr/UYZS
[20] Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 3645-3650). https://doi.org/10.18653/v1/P19-1355