Important notice
The course guide is provisional.
The PDF version of the course guide may take a few days to become available in the DDD.

3D Vision
Code: 44775Credits: 6
| Degree programme | Type | Course |
|---|---|---|
| Computer Vision | OB | 1 |
Contact lecturer
- Name :
- Maria Vanrell Martorell
- Email :
- maria.vanrell@uab.cat
Teaching staff
- Josep Ramon Casas Pla
- Javier Ruiz Hidalgo
- Gloria Haro Ortega
- Antonio Agudo Martínez
- Marc Pérez Quintana
- Federico Sukno
Group languages
You can consult this information at the end of the document.
Prerequisites
Degree in Engineering, Maths, Physics or similar
Objectives
Module Coordinator: Dr. Gloria Haro
The goal of this module is to learn the principles of the 3D reconstruction of an object or a scene from multiple images or stereoscopic videos. For that, the basic concepts of the projective geometry and the 3D space are firstly introduced. The rest of the theoretical aspects and applications are built upon these basic tools. The mapping from the 3D world to the image plane will be studied, for that we will introduce different camera models, their parameters and how to estimate them (camera calibration and auto-calibration). The geometry that relates a pair of views will be analyzed. All these concepts will be applied to obtain a 3D reconstruction in the two main possible settings: calibrated or uncalibrated cameras. In particular, we will learn how to: estimate the depth of image points, extract the underlying 3D points given a set of point correspondences in the images, generate novel views, estimate the 3D object given a set of calibrated color images or binary images, and estimate a sparse set of 3D points given a set of uncalibrated images. The 3D representation in voxels and meshes will be studied. We will explain the reconstruction and modeling from Kinect data, as a particular model of sensors that provide an image of the scene together with its depths. Finally, we will see some techniques for processing 3D point clouds. The concepts and techniques learnt in this module are used in real applications ranging from augmented reality, object scanning, motion capture, new view synthesis, bullet-time effect, robotics, etc.
Learning outcomes
- CA01 (Integrate the formulation of all the components of a complete system of 3D information recovery and synthesis.) Integrate the formulation of all the components of a complete system of 3D information recovery and synthesis.
- CA06 (Achieve the objectives of a project of vision carried out in a team.) Achieve the objectives of a project of vision carried out in a team.
- KA04 (Identify the basic problems to be solved in a case of 3D information recovery from a scene.) Identify the basic problems to be solved in a case of 3D information recovery from a scene.
- KA12 (Provide the best geometric information needed to model all parts of a problem of recovering 3D information from a scene.) Provide the best geometric information needed to model all parts of a problem of recovering 3D information from a scene.
- SA04 (Solve a problem of 3D information recovery and evaluate the results.) Solve a problem of 3D information recovery and evaluate the results.
- SA10 (Define the best data sets for training 3D vision architectures.) Define the best data sets for training 3D vision architectures.
- SA15 (Prepare a report that describes, justifies and illustrates the development of a project of vision.) Prepare a report that describes, justifies and illustrates the development of a project of vision.
- SA17 (Prepare oral presentations that allow debate of the results of a project of vision.) Prepare oral presentations that allow debate of the results of a project of vision.
Contents
- Introduction and applications.
- 2D projective geometry. Planar transformations.
- Homography estimation. Affine and metric rectification
- 3D projective geometry and transformations. Camera models
- Camera calibration. Pose estimation.
- Epipolar geometry. Fundamental matrix. Essential matrix. Extraction of camera matrices.
- Computation of the fundamental matrix. Image rectification
- Triangulation methods. Depth computation. New View Synthesis
- Multi-view stereo. Structure from motion.
- Auto-calibration. Bundle adjustment.
- 3D sensors (kinect).
- Point cloud processing.
Learning activities and methodology
| Title | Hours | ECTS | Learning outcomes |
|---|---|---|---|
| Project follow-up sessions | 8 | 0.32 | CA01, CA06, KA04, KA12, SA04, SA10, SA15, SA17 |
| Lecture sessions | 20 | 0.8 | CA01, CA06, KA04, KA12, SA04, SA10, SA15, SA17 |
| Homework | 113 | 4.52 | CA01, CA06, KA04, KA12, SA04, SA10, SA15, SA17 |
Supervised sessions: (Some of these sessions could be synchronous on-line sessions)
- Lecture Sessions, where the lecturers will explain general contents about the topics. Some of them will be used to solve the problems.
Directed sessions:
- Project Sessions, where the problems and goals of the projects will be presented and discussed, students will interact with the project coordinator about problems and ideas on solving the project (approx. 1 hour/week)
- Presentation Session, where the students give an oral presentation about how they have solved the project and a demo of the results.
- Exam Session, where the students are evaluated individually. Knowledge achievements and problem-solving skills
Autonomous work:
- Student will autonomously study and work with the materials derived from the lectures.
- Student will work in groups to solve the problems of the projects with deliverables:
- Code
- Reports
- Oral presentations
Assessment
Continuous assessment activities
| Title | Weight | Hours | ECTS | Learning outcomes |
|---|---|---|---|---|
| Project | 0.55 | 6 | 0.24 | CA01, CA06, KA04, KA12, SA04, SA10, SA15, SA17 |
| Session attendance | 0.05 | 0.5 | 0.02 | CA01, CA06, KA04, KA12, SA04, SA10, SA15, SA17 |
| Exam | 0.4 | 2.5 | 0.1 | CA01, KA04, KA12, SA04, SA10 |
The final mark for this module will be computed with the following formula:
Final Mark = 0.4 x Exam + 0.55 x Project+ 0.05 x Attendance
where,
- Exam: is the mark obtained in the Module Exam (must be >= 4).
- Attendance: is the mark derived from the control of attendance at lectures (minimum 70%)
- Projects: is the mark provided by the project coordinator based on the weekly follow-up of the project and deliverables (must be >= 5). All accordingly with specific criteria such as:
- Participation in discussion sessions and in team work (inter-member evaluations)
- Delivery of mandatory and optional exercises.
- Code development (style, comments, etc.)
- Report (justification of the decisions in your project development)
- Presentation and exam (Talk, explanations and demonstrations on your project)
Only those students that fail (Final Mark < 5.0) can do a retake exam.
Bibliography
Books:
- O. Faugeras, Three-dimensional computer vision: a geometric viewpoint, MIT Press, cop. 1993.
- O. Faugeras, Q.T. Loung, The geometry of multiple images, MIT Press, 2001.
- D. A. Forsyth, J. Ponce, Computer vision: a modern approach, Prentice Hall, 2003.
- R. I. Hartley, A. Zisserman, Multiple view geometry in computer vision, Cambridge University Press, 2000.
- R. Szeliski, Computer Vision: Algorithms and Applications, Springer, 2011.
Tutorials:
- Y. Furukawa and C. Hernández, Multi-View Stereo: A Tutorial, Foundations and Trends® in Computer Graphics and Vision, vol. 9, no. 1-2, pp.1-148, 2013.
- T. Moons, L. Van Gool, M. Vergauwen, 3D Reconstruction from Multiple Images Part 1, Principles, Foundations and Trends® in Computer Graphics and Vision, vol. 4: no. 4, pp 287-404, 2010.
Software
Python Programming Tools with special attention to computer vision and image processing libraries
Course groups and languages
The information provided is provisional until November 30. After this date, you will be able to consult the language of each group through this link. To access the information, you will need to enter the course CODE