Logo

Advanced Artificial Vision

Code: 108269
Credits: 6
2026/2027
Degree programme Type Course
Bachelor in Artificial Intelligence OP 3

Contact lecturer

Name :
Alexandra Gomez Villa
Email :
alexandra.gomez@uab.cat

Group languages

You can consult this information at the end of the document.

Prerequisites

Recommended "Fundamentals of Machine Learning", "Neural Networks and Deep Learning", and “Fundamentals of Computer Vision”


Objectives

This subject aims to provide students with a comprehensive understanding of advanced artificial vision, covering both geometric 3D vision and modern deep generative modeling. The course is organized into two complementary blocks. The first half builds on fundamental computer vision concepts to explore how 3D scenes are represented, reconstructed, and rendered: students will work with classical 3D representations, understand image formation and camera geometry through structure from motion, and study how deep learning has transformed novel view synthesis through implicit neural representations, Neural Radiance Fields (NeRF), and Gaussian Splatting. The second half shifts focus to deep generative models, providing students with the theoretical and practical foundations of Variational Autoencoders (VAEs), flow-based models, diffusion models, and flow matching. Throughout the course, connections between the two blocks will be emphasized, particularly how generative modeling techniques increasingly intersect with 3D vision in tasks such as novel view synthesis, 3D asset generation, and scene reconstruction. By the end of this subject, students should be able to design and implement learning-based solutions for both 3D vision and generative modeling problems, understand the theoretical underpinnings and practical trade-offs of current approaches, and apply these techniques in practical applications ranging from robotics and autonomous driving to virtual reality, computer graphics, and content generation.

Learning outcomes

  • CM12 (Build solutions to complex artificial intelligence problems based on the selection of the most appropriate image processing and computational learning techniques applied to computer vision.) Build solutions to complex artificial intelligence problems based on the selection of the most appropriate image processing and computational learning techniques applied to computer vision.
  • CM13 (Plan the development and deployment of solutions based on computer vision using the appropriate tools and platforms, taking into account sustainability criteria in the use of computational resources.) Plan the development and deployment of solutions based on computer vision using the appropriate tools and platforms, taking into account sustainability criteria in the use of computational resources.
  • KM29 (Identify the mathematical and algorithmic underpinnings on which low-level image processing, optimisation, and computational learning techniques applied to computer vision are based.) Identify the mathematical and algorithmic underpinnings on which low-level image processing, optimisation, and computational learning techniques applied to computer vision are based.
  • KM30 (Identify the most common deep learning architectures and models for solving computer vision problems) Identify the most common deep learning architectures and models for solving computer vision problems
  • SM31 (Apply the right machine learning techniques to machine vision problems.) Apply the right machine learning techniques to machine vision problems.
  • SM33 (Perform data preparation, training, and model validation and analysis on all types of machine vision problems.) Perform data preparation, training, and model validation and analysis on all types of machine vision problems.

Contents

Block 1: 3D Vision


3D Representations

  • Classical 3D representations: point clouds, meshes, voxels
  • Transformations between 3D representations


Structure from Motion

  • Camera models and calibration
  • Feature matching and geometric verification
  • Triangulation


Neural Rendering

  • Implicit neural representations
  • Neural Radiance Fields (NeRF)
  • View synthesis and novel view generation


Gaussian Splatting

  • 3D Gaussian primitives
  • Differentiable rasterization
  • Applications in neural rendering


Block 2: Deep Generative Modeling


Variational Autoencoders

  • Latent variable models
  • Encoder-decoder architectures and the reparameterization trick
  • Evidence lower bound (ELBO) and variational inference
  • Applications in generative modeling


Flow-based Models

  • Change of variables and normalizing flows
  • Coupling layers and invertible architectures
  • Density estimation and exact likelihood computation
  • Applications in generative modeling


Diffusion Models

  • Foundations of diffusion models
  • Forward and reverse diffusion processes
  • Conditioning strategies
  • Applications in image and 3D generation


Flow Matching

  • Flow matching objective and training
  • Connections to diffusion models
  • Applications in generative modeling

Learning activities and methodology

Title Hours ECTS Learning outcomes
Work in projects/exercises 42 1.68 CM13, KM30, SM31, SM33
Individual study 20 0.8 CM12, KM29, KM30, SM31
Laboratory sessions 21 0.84 CM12, CM13, KM29, KM30, SM31, SM33
Theory classes 24 0.96 CM12, CM13, KM29, KM30, SM33
Work in projects 30 1.2 CM12, CM13, KM30, SM31, SM33

Real-world perception and content-generation challenges guide this course. Throughout the subject, practical applications in robotics, autonomous driving, virtual reality, computer graphics, and generative content creation will motivate each section and direct the organization of the content.


There will be two types of sessions:


Theory classes: The objective of these sessions is for the teacher to explain the theoretical foundations of 3D vision and deep generative modeling. For each topic studied, the underlying mathematical principles — of 3D geometry in the first half, and of probabilistic and generative modeling in the second half — will be explained, as well as the corresponding algorithmic implementations. Topics will range from classical 3D representations and structure-from-motion to neural rendering and Gaussian Splatting, followed by the theoretical foundations of generative modeling, including VAEs, flow-based models, diffusion models, and flow matching.


Laboratory sessions: Laboratory sessions aim to facilitate hands-on experience with 3D vision and generative modeling systems, reinforcing the concepts covered in theory classes. During the first half, students will work through practical cases implementing solutions using PyTorch3D and other relevant 3D vision libraries, gaining experience with tasks such as camera calibration, neural rendering, and Gaussian Splatting. During the second half, students will implement generative models in Python (using PyTorch), including training VAEs, normalizing flows, and diffusion models, as well as flow-matching approaches, on practical generation tasks. The sessions will emphasize collaborative work and provide students with direct experience applying theoretical concepts to real problems. Problem-solving will be initiated in class and complemented by a weekly set of problems to work through at home.


Course Project: A course project will be carried out during the semester, in which students will tackle a challenging problem spanning 3D vision, generative modeling, or their intersection. Examples of projects might include developing neural rendering or Gaussian Splatting pipelines, training generative models for image or 3D asset synthesis, or combining generative modeling techniques with 3D representations for tasks such as 3D-aware generation or scene reconstruction. Students will work in small groups of 2-3, where each member must equally contribute to the final solution. These working groups will be maintained until the end of the semester. They must be self-managed in terms of role distribution, work planning, task assignment, resource management, conflict resolution, etc. To develop the project, the groups will work autonomously, while some of the laboratory sessions will be used (1) for the teacher to present the project's theme and discuss possible approaches, and (2) for monitoring the status of the project.


All information on the subject and related documents students need will be available on the virtual campus (cv.uab.cat).


Annotation: Within the schedule set by the center or degree program, 15 minutes of one class will be reserved for students to evaluate their lecturers and their courses or modules through questionnaires.


Annotation: within the schedule set by the centre or degree programme, 15 minutes of one class will be reserved for students to evaluate their lecturers and their courses or modules through questionnaires.

Assessment

Continuous assessment activities

Title Weight Hours ECTS Learning outcomes
Written exams 0.4 4 0.16 CM12, CM13, KM29, KM30, SM31
Delivery of exercises 0.1 5 0.2 CM12, CM13, KM29, KM30, SM33
Delivery of project 0.5 4 0.16 CM12, CM13, KM30, SM31, SM33

The final grade assessment combines three components to evaluate theoretical knowledge, practical application, and problem-solving abilities.

Final Grade Calculation

The final grade is calculated using the following weighted formula:

Final Grade = (40% × Theory Grade) + (10% × Problems Portfolio Grade) + (50% × Project Grade)

To pass the course, students must achieve a minimum grade of 5.0 in both the Theory Grade and Project Grade components. While there is no minimum grade requirement for the Problems Portfolio, it's important to note a special condition: if a student's weighted calculation results in a grade of 5.0 or higher, but either their Theory Grade or Project Grade falls below 5.0, their final grade will be automatically capped at 4.5, regardless of their overall weighted average.

Theory Grade Assessment

The Theory Grade is designed to assess each student's individual mastery of course content through a continuous assessment model utilizing two examinations. The first is the Mid-term Examination (Exam 1), which takes place mid-semester and covers the first half of the course material. The second is the Final Examination (Exam 2), which is conducted at the end of the semester and focuses on the second half of the course materials. The Theory Grade is determined by calculating the average of these two examinations.

Theory Grade = (Exam 1 + Exam 2) ÷ 2

The examinations are structured to evaluate two critical components: students' problem-solving abilities using techniques covered in class, and their conceptual understanding of these techniques. There are specific requirements for the Theory Grade: students must score above 4.0 on both partial exams. If a student achieves an average of 5.0 or higher across both exams but scores below 4.0 on either exam, their theory grade will be adjusted down to 4.5 for the final grade calculation. For students who do not achieve a passing theory grade, there is an opportunity to take a recovery examination, which allows them to retake the failed portions (either Part 1, Part 2, or both) of the continuous evaluation process.

Problems Portfolio Assessment

The problems portfolio serves a dual purpose: to promote continuous engagement with course material and to provide opportunities for the practical application of theoretical concepts. To successfully complete the portfolio, students must submit at least 70% of all assigned problem sets. It's important to note that failing to meet this 70% submission threshold will automatically result in a Problems Portfolio Grade of Zero. Students should compile and document all their completed exercises as part of the portfolio requirements.

Project Assessment

The project stands as a core component of the course, designed to fulfill multiple educational objectives. Students are required to work collaboratively in teams, develop a comprehensive solution to a given challenge, demonstrate their teamwork capabilities, and present their findings to the class. The project grade is determined through a calculation that takes into account various evaluation components and their respective weightings.


Project Grade = (80% × Deliverables Grade) + (20% × Presentation Grade)


The project has several mandatory requirements: students must participate in all three components (deliverables, presentation, and self-evaluation). If a student achieves a weighted calculation of 5.0 or higher but fails to participate in any component, their project grade will be adjusted down to 4.5. While students have the opportunity to resubmit failed projects for recovery, these grades will be capped at 7.0 out of 10. There are also several conditions that result in automatic failure of the project: these include non-submission of project deliverables, failure to present the project, use of copied content, or use of synthetically generated content.

Important notes

Academic Misconduct Any academic irregularities, including plagiarism, copying, or allowing copying, will result in an immediate grade of zero for the relevant assessment activity. These activities cannot be recovered if failed due to academic misconduct. Moreover, if passing the failed activity is required to pass the course, the student will fail the entire course without any opportunity for recovery within the same academic year.


Grade Assignment

A grade of "non-evaluable" will be assigned only in specific circumstances: when a student submits no exercise solutions, does not participate in any practical evaluations, and takes no examinations. In all other situations, missing submissions or non-participation will be counted as a zero in the weighted average calculation.


Honors Qualification (Distinction)

Students may qualify for honors distinction by achieving a final grade of 9.0 or higher. However, this distinction is limited to 5% of enrolled students and is awarded only to those with the highest final marks. In situations where there are ties among students, partial exam results will be used as the deciding factor.


Policy Updates

The detailed evaluation process will be explained during the first weeks of the semester. It's important to note that if any discrepancies exist between this guide and the information provided in class, the information provided in class will take precedence.


Bibliography

Block 1: 3D Vision

  • Hartley, R., & Zisserman, A. (2003). Multiple view geometry in computer vision. Cambridge University Press.
  • Pharr, M., Jakob, W., & Humphreys, G. (2023). Physically based rendering: From theory to implementation. MIT Press.
  • Malhotra, A., & Seghouani, N. (2025). Gaussian Splatting: An Introduction


Block 2: Deep Generative Modeling

  • Tomczak, J. M. (2024). Deep generative modeling (2nd ed.). Springer.
  • Lai, C.-H., Song, Y., Kim, D., Mitsufuji, Y., & Ermon, S. (2025). The principles of diffusion models. arXiv:2510.21890. Website: https://the-principles-of-diffusion-models.github.io (Forthcoming, MIT Press, 2027.)
  • Lipman, Y., Havasi, M., Holderrieth, P., Shaul, N., Le, M., Karrer, B., Chen, R. T. Q., Lopez-Paz, D., Ben-Hamu, H., & Gat, I. (2024). Flow matching guide and code. arXiv:2412.06264.

Software

  • VScode

Course groups and languages

The information provided is provisional until November 30. After this date, you will be able to consult the language of each group through this link. To access the information, you will need to enter the course CODE

Type of teaching Group Language Semester Shift
(TE) Theory 1 English second semester afternoon
(PAUL) Classroom practices 1 English second semester afternoon