The explanation of "What's that extra 1 for?" in the column representation of 3-d coordinates (x y z 1) could benefit from mentioning that translation—moving things—is not a linear transformation (the origin is not mapped to itself) but an affine transformation. Therefore, you cannot represent translation in 3-d space with a 3x3 matrix. What you can do, though, is embed that 3-d space within a 4-d space fixed at some coordinate on its 4th dimension, typically w=1. Then, a translation in the original 3-d space can be represented as a linear transformation in the 4-d space and thus can also be represented by a 4x4 matrix multiplication. So the extra 1 is actually what allows all common 3-d operations, including translation, to be done via linear algebra and thereby harness the brutal power of matrix multiplication on modern computing devices.
You can do rotations and translations "via linear algebra" without homogeneous coordinates (the 4-element tuple representing a 3d point or vector), since adding two vectors is linear algebra.
What this lets you do is composing multiple transforms into one matrix multiplication instead of a sequence of multiplications and additions; that's what dramatically increases performance, on modern computing devices but most especially on ancient ones, where we were fighting for every MUL.
More details: https://gabrielgambetta.com/computer-graphics-from-scratch/1...
You don't need to use homogeneous coordinates to represent a series of linear transformations as a single matrix. That's just basic linear algebra: any series of linear transformations is itself a linear transformation, and any linear transformation can be represented as an equivalent transformation matrix.
But you do need homogenous coordinates if you want to include translations (fixed distance shifts, such as moving the camera) in the set of operations you can represent as linear transformations and thereby gain all the benefits of linear algebra, including the benefit of being able to include them in a series of operations that you can represent with a single transformation matrix.
Rick Szeliaki’s “Computer vision algorithms and applications” has a lot of great motivation for homogeneous coordinates in the first chapter and appendix. For example, in 2D, the cross product between two points in homogeneous coordinates is the line joining them, and the cross product between two lines in homogeneous coordinates is their point of intersection.
its linear algebra but not a linear transformation. repeat: translations are not linear transformations in 3d.
Is that why quaternions are interesting for computer graphics? I studied mathematics and computer science but never understood (or even looked so hard into) the connection.
Kind of! The scalar allows vectors and scalars to be expressed together as one object, which has a lot of computational and mathematical niceties.
The killer feature is that you can put the two together for a rotation axis (~ vector) and angle (~ scalar).
With just three gimbals (rotating circles), if gimbals A and B are aligned, you only have two degrees of freedom (rotating A is the same as rotating B). Because of this, interpolating angles is unwieldy in vector space. 'Gimbal lock' confounds animation (in hilarious but unrealistic ways) but also aerospace (four hours before 'one small step for man', just after landing, Collins joked he would like a fourth gimbal for Christmas).
nice intuition there, this comment prompted me to consider a simpler example, 2d embedded within 3d. does the 2d plane (embedded in 3d) go through the 3d origin (where 0 maps to 0) and is thus a linear transformation in 3d but a 2d affine transform in 2d? it feels like this is the case?
No, the 2d x-y plane in your example cannot pass through the 3d space’s origin because that would imply that you fixed the z coordinate at zero. The plane must be fixed at some nonzero z because you need to be able move x and y values by some scaled version of z to make translation happen. If z is zero, that scheme does not work.
Consider a transformation f where we wish to move x-y coordinates s units to the right. In 2d, we could express it as:
f(x, y) = (x + s, y)
But that transformation is affine not linear. There is no way to generate the value s as a linear combination of the inputs x and y. So, our workaround is to embed the x-y plane into 3d space at z=1. Then we can move (x,y,1) points in that plane s units to the right using this transformation:
f(x, y, z) = (x + s*z, y, z)
This new transformation is linear: it maps (0,0,0) to itself. But it maps our embedded 2d plane's origin (0,0,1) to (s,0,1), shifting it right by s units, as we want.
The matrix form of that transformation is:
The same scheme would work if we had embedded the plane at any fixed z=r for nonzero r. We would only have to rescale the s in the matrix to s/r. Again, however, if r=0, this scheme will not work, as 1/r has gone to infinity.Yes, affine transformation matrices are essentially shears.[0] In the 2D case, shearing the plane z=1 in 3D space essentially translates it around.
[0]: Here’s a visual: https://gunn-gatm.github.io/textbook/gatm.pdf#page=28
ive been thinking -- if i have a 3x3 matrix [1 0 q; 0 1 r; 0 0 1], when dotted with x, those rows are planes in 3d with normals n1=(1 0 q), n2=(0 1 r) and n3=(0 0 1). plane 3 is parallel to the x-y plane and 1 unit up. plane 1 is tilted by q and parallel to the y-axis, plane 2 is tilted by r and parallel to the x-axis. the intersection of those planes (i.e. the solution x) when calculating Ax=b gives an output vector b sitting in 3d space at b=(x+qz, y+rz, 1*z). since we always specify z=1 we have b=(x+q, y+r, 1). in 3d this is a linear shear because we are translating proportionally by z but because z always equals 1 in this case we effectively get a translation in 2d.
so as the other 2 helpful commenters also just said: 3d shears using linear algebra degenerate to 2d affine transformations when z=1 (or w in 4d)