
Visió Artificial Avançada
Codi: 108269Crèdits: 6
| Titulació | Tipus | Curs |
|---|---|---|
| Intel·ligència Artificial / Bachelor in Artificial Intelligence | OP | 3 |
Professor/a de contacte
- Nom :
- Alexandra Gomez Villa
- Correu electrònic :
- alexandra.gomez@uab.cat
Idiomes dels grups
Podeu consultar aquesta informació al final del document.
Prerequisits
Recommended "Fundamentals of Machine Learning", "Neural Networks and Deep Learning", and “Fundamentals of Computer Vision”
Objectius
This subject aims to provide students with a comprehensive understanding of advanced artificial vision, covering both geometric 3D vision and modern deep generative modeling. The course is organized into two complementary blocks. The first half builds on fundamental computer vision concepts to explore how 3D scenes are represented, reconstructed, and rendered: students will work with classical 3D representations, understand image formation and camera geometry through structure from motion, and study how deep learning has transformed novel view synthesis through implicit neural representations, Neural Radiance Fields (NeRF), and Gaussian Splatting. The second half shifts focus to deep generative models, providing students with the theoretical and practical foundations of Variational Autoencoders (VAEs), flow-based models, diffusion models, and flow matching. Throughout the course, connections between the two blocks will be emphasized, particularly how generative modeling techniques increasingly intersect with 3D vision in tasks such as novel view synthesis, 3D asset generation, and scene reconstruction. By the end of this subject, students should be able to design and implement learning-based solutions for both 3D vision and generative modeling problems, understand the theoretical underpinnings and practical trade-offs of current approaches, and apply these techniques in practical applications ranging from robotics and autonomous driving to virtual reality, computer graphics, and content generation.
Resultats d'aprenentatge
- CM12 (Construir solucions a problemes complexos d’intel·ligència artificial basades en la selecció de les tècniques més adequades de processament d’imatges i aprenentatge computacional aplicat a la visió artificial.) Construir solucions a problemes complexos d’intel·ligència artificial basades en la selecció de les tècniques més adequades de processament d’imatges i aprenentatge computacional aplicat a la visió artificial.
- CM13 (Planificar el desenvolupament i desplegament de solucions basades en visió per computador utilitzant les eines i plataformes adequades, tenint en compte criteris de sostenibilitat en l’ús dels recursos computacionals.) Planificar el desenvolupament i desplegament de solucions basades en visió per computador utilitzant les eines i plataformes adequades, tenint en compte criteris de sostenibilitat en l’ús dels recursos computacionals.
- KM29 (Identificar els fonaments matemàtics i algorítmics en els quals es basen les tècniques de processament d’imatges de baix nivell, d’optimització i d’aprenentatge computacional aplicat a la visió artificial.) Identificar els fonaments matemàtics i algorítmics en els quals es basen les tècniques de processament d’imatges de baix nivell, d’optimització i d’aprenentatge computacional aplicat a la visió artificial.
- KM30 (Identificar les arquitectures i els models d’aprenentatge profund més comuns per a la resolució de problemes de visió artificial.) Identificar les arquitectures i els models d’aprenentatge profund més comuns per a la resolució de problemes de visió artificial.
- SM31 (Aplicar les tècniques d’aprenentatge automàtic adequades a problemes de visió artificial.) Aplicar les tècniques d’aprenentatge automàtic adequades a problemes de visió artificial.
- SM33 (Dur a terme la preparació de dades, l’entrenament i la validació i l’anàlisi de models en tot tipus de problemes de visió artificial.) Dur a terme la preparació de dades, l’entrenament i la validació i l’anàlisi de models en tot tipus de problemes de visió artificial.
Continguts
Block 1: 3D Vision
3D Representations
- Classical 3D representations: point clouds, meshes, voxels
- Transformations between 3D representations
Structure from Motion
- Camera models and calibration
- Feature matching and geometric verification
- Triangulation
Neural Rendering
- Implicit neural representations
- Neural Radiance Fields (NeRF)
- View synthesis and novel view generation
Gaussian Splatting
- 3D Gaussian primitives
- Differentiable rasterization
- Applications in neural rendering
Block 2: Deep Generative Modeling
Variational Autoencoders
- Latent variable models
- Encoder-decoder architectures and the reparameterization trick
- Evidence lower bound (ELBO) and variational inference
- Applications in generative modeling
Flow-based Models
- Change of variables and normalizing flows
- Coupling layers and invertible architectures
- Density estimation and exact likelihood computation
- Applications in generative modeling
Diffusion Models
- Foundations of diffusion models
- Forward and reverse diffusion processes
- Conditioning strategies
- Applications in image and 3D generation
Flow Matching
- Flow matching objective and training
- Connections to diffusion models
- Applications in generative modeling
Activitats formatives i Metodologia
| Títol | Hores | ECTS | Resultats d'aprenentatge |
|---|---|---|---|
| Work in projects/exercises | 42 | 1,68 | CM13, KM30, SM31, SM33 |
| Individual study | 20 | 0,8 | CM12, KM29, KM30, SM31 |
| Laboratory sessions | 21 | 0,84 | CM12, CM13, KM29, KM30, SM31, SM33 |
| Theory classes | 24 | 0,96 | CM12, CM13, KM29, KM30, SM33 |
| Work in projects | 30 | 1,2 | CM12, CM13, KM30, SM31, SM33 |
Real-world perception and content-generation challenges guide this course. Throughout the subject, practical applications in robotics, autonomous driving, virtual reality, computer graphics, and generative content creation will motivate each section and direct the organization of the content.
There will be two types of sessions:
Theory classes: The objective of these sessions is for the teacher to explain the theoretical foundations of 3D vision and deep generative modeling. For each topic studied, the underlying mathematical principles — of 3D geometry in the first half, and of probabilistic and generative modeling in the second half — will be explained, as well as the corresponding algorithmic implementations. Topics will range from classical 3D representations and structure-from-motion to neural rendering and Gaussian Splatting, followed by the theoretical foundations of generative modeling, including VAEs, flow-based models, diffusion models, and flow matching.
Laboratory sessions: Laboratory sessions aim to facilitate hands-on experience with 3D vision and generative modeling systems, reinforcing the concepts covered in theory classes. During the first half, students will work through practical cases implementing solutions using PyTorch3D and other relevant 3D vision libraries, gaining experience with tasks such as camera calibration, neural rendering, and Gaussian Splatting. During the second half, students will implement generative models in Python (using PyTorch), including training VAEs, normalizing flows, and diffusion models, as well as flow-matching approaches, on practical generation tasks. The sessions will emphasize collaborative work and provide students with direct experience applying theoretical concepts to real problems. Problem-solving will be initiated in class and complemented by a weekly set of problems to work through at home.
Course Project: A course project will be carried out during the semester, in which students will tackle a challenging problem spanning 3D vision, generative modeling, or their intersection. Examples of projects might include developing neural rendering or Gaussian Splatting pipelines, training generative models for image or 3D asset synthesis, or combining generative modeling techniques with 3D representations for tasks such as 3D-aware generation or scene reconstruction. Students will work in small groups of 2-3, where each member must equally contribute to the final solution. These working groups will be maintained until the end of the semester. They must be self-managed in terms of role distribution, work planning, task assignment, resource management, conflict resolution, etc. To develop the project, the groups will work autonomously, while some of the laboratory sessions will be used (1) for the teacher to present the project's theme and discuss possible approaches, and (2) for monitoring the status of the project.
All information on the subject and related documents students need will be available on the virtual campus (cv.uab.cat).
Annotation: Within the schedule set by the center or degree program, 15 minutes of one class will be reserved for students to evaluate their lecturers and their courses or modules through questionnaires.
Avaluació
Activitats d'avaluació continuada
| Títol | Pes | Hores | ECTS | Resultats d'aprenentatge |
|---|---|---|---|---|
| Written exams | 0.4 | 4 | 0,16 | CM12, CM13, KM29, KM30, SM31 |
| Delivery of exercises | 0.1 | 5 | 0,2 | CM12, CM13, KM29, KM30, SM33 |
| Delivery of project | 0.5 | 4 | 0,16 | CM12, CM13, KM30, SM31, SM33 |
The final grade assessment combines three components to evaluate theoretical knowledge, practical application, and problem-solving abilities.
Final Grade Calculation
The final grade is calculated using the following weighted formula:
Final Grade = (40% × Theory Grade) + (10% × Problems Portfolio Grade) + (50% × Project Grade)
To pass the course, students must achieve a minimum grade of 5.0 in both the Theory Grade and Project Grade components. While there is no minimum grade requirement for the Problems Portfolio, it's important to note a special condition: if a student's weighted calculation results in a grade of 5.0 or higher, but either their Theory Grade or Project Grade falls below 5.0, their final grade will be automatically capped at 4.5, regardless of their overall weighted average.
Theory Grade Assessment
The Theory Grade is designed to assess each student's individual mastery of course content through a continuous assessment model utilizing two examinations. The first is the Mid-term Examination (Exam 1), which takes place mid-semester and covers the first half of the course material. The second is the Final Examination (Exam 2), which is conducted at the end of the semester and focuses on the second half of the course materials. The Theory Grade is determined by calculating the average of these two examinations.
Theory Grade = (Exam 1 + Exam 2) ÷ 2
The examinations are structured to evaluate two critical components: students' problem-solving abilities using techniques covered in class, and their conceptual understanding of these techniques. There are specific requirements for the Theory Grade: students must score above 4.0 on both partial exams. If a student achieves an average of 5.0 or higher across both exams but scores below 4.0 on either exam, their theory grade will be adjusted down to 4.5 for the final grade calculation. For students who do not achieve a passing theory grade, there is an opportunity to take a recovery examination, which allows them to retake the failed portions (either Part 1, Part 2, or both) of the continuous evaluation process.
Problems Portfolio Assessment
The problems portfolio serves a dual purpose: to promote continuous engagement with course material and to provide opportunities for the practical application of theoretical concepts. To successfully complete the portfolio, students must submit at least 70% of all assigned problem sets. It's important to note that failing to meet this 70% submission threshold will automatically result in a Problems Portfolio Grade of Zero. Students should compile and document all their completed exercises as part of the portfolio requirements.
Project Assessment
The project stands as a core component of the course, designed to fulfill multiple educational objectives. Students are required to work collaboratively in teams, develop a comprehensive solution to a given challenge, demonstrate their teamwork capabilities, and present their findings to the class. The project grade is determined through a calculation that takes into account various evaluation components and their respective weightings.
Project Grade = (80% × Deliverables Grade) + (20% × Presentation Grade)
The project has several mandatory requirements: students must participate in all three components (deliverables, presentation, and self-evaluation). If a student achieves a weighted calculation of 5.0 or higher but fails to participate in any component, their project grade will be adjusted down to 4.5. While students have the opportunity to resubmit failed projects for recovery, these grades will be capped at 7.0 out of 10. There are also several conditions that result in automatic failure of the project: these include non-submission of project deliverables, failure to present the project, use of copied content, or use of synthetically generated content.
Important notes
Academic Misconduct Any academic irregularities, including plagiarism, copying, or allowing copying, will result in an immediate grade of zero for the relevant assessment activity. These activities cannot be recovered if failed due to academic misconduct. Moreover, if passing the failed activity is required to pass the course, the student will fail the entire course without any opportunity for recovery within the same academic year.
Grade Assignment
A grade of "non-evaluable" will be assigned only in specific circumstances: when a student submits no exercise solutions, does not participate in any practical evaluations, and takes no examinations. In all other situations, missing submissions or non-participation will be counted as a zero in the weighted average calculation.
Honors Qualification (Distinction)
Students may qualify for honors distinction by achieving a final grade of 9.0 or higher. However, this distinction is limited to 5% of enrolled students and is awarded only to those with the highest final marks. In situations where there are ties among students, partial exam results will be used as the deciding factor.
Policy Updates
The detailed evaluation process will be explained during the first weeks of the semester. It's important to note that if any discrepancies exist between this guide and the information provided in class, the information provided in class will take precedence.
Bibliografia
Block 1: 3D Vision
- Hartley, R., & Zisserman, A. (2003). Multiple view geometry in computer vision. Cambridge University Press.
- Pharr, M., Jakob, W., & Humphreys, G. (2023). Physically based rendering: From theory to implementation. MIT Press.
- Malhotra, A., & Seghouani, N. (2025). Gaussian Splatting: An Introduction
Block 2: Deep Generative Modeling
- Tomczak, J. M. (2024). Deep generative modeling (2nd ed.). Springer.
- Lai, C.-H., Song, Y., Kim, D., Mitsufuji, Y., & Ermon, S. (2025). The principles of diffusion models. arXiv:2510.21890. Website: https://the-principles-of-diffusion-models.github.io (Forthcoming, MIT Press, 2027.)
- Lipman, Y., Havasi, M., Holderrieth, P., Shaul, N., Le, M., Karrer, B., Chen, R. T. Q., Lopez-Paz, D., Ben-Hamu, H., & Gat, I. (2024). Flow matching guide and code. arXiv:2412.06264.
Programari
VScode
Grups i idiomes de l'assignatura
La informació proporcionada és provisional fins al 30 de novembre. A partir d'aquesta data, podreu consultar l'idioma de cada grup a través d'aquest enllaç. Per accedir a la informació, caldrà introduir el CODI de l'assignatura
| Tipus de docència | Grup | Idioma | Semestre | Torn |
|---|---|---|---|---|
| (TE) Teoria | 1 | Anglès | segon quadrimestre | tarda |
| (PAUL) Pràctiques d'aula | 1 | Anglès | segon quadrimestre | tarda |