Artificial Intelligence Applications in Engineering (MUH-920), week 13 of 14: interactive lab

Language lab: Tokens, sampling and grounded retrieval

Prof. Dr. Utku Kose, Süleyman Demirel University

Part A learns byte-pair merges from engineering sentences and shows how any typed text splits into tokens [1]. Part B turns next-token scores into probabilities with an adjustable temperature and top-k cut, and samples continuations. Part C retrieves passages from a small engineering knowledge base with TF-IDF and assembles an answer that cites them, as retrieval-augmented generation does [2, 3].

Part A: Byte-pair tokenisation

Part B: Temperature and top-k sampling

Context: "The concrete cubes were tested after twenty eight ..."

Part C: Retrieval with citations

References

[1] Sennrich, R., Haddow, B., & Birch, A. (2016). Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (pp. 1715-1725). https://doi.org/10.18653/v1/P16-1162

[2] Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press.

[3] Lewis, P., Perez, E., Piktus, A., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems 33 (pp. 9459-9474). https://arxiv.org/abs/2005.11401

[4] Ouyang, L., Wu, J., Jiang, X., et al. (2022). Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems 35 (pp. 27730-27744). https://arxiv.org/abs/2203.02155

[5] Ji, Z., Lee, N., Frieske, R., et al. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 248. https://doi.org/10.1145/3571730