Defining Our Goal

Suppose you place two points on the x-y plane and ask a seemingly simple question: What is the shortest path connecting them?

The "obvious" answer is a straight line (the green curve), but how can we show that this is truly the optimal path? Imagine two fixed points (x1,y1) and (x2,y2) connected by some curve y(x).

To find the length of this curve, we zoom in on an infinitesimally small segment. If the curve changes by small amounts dx horizontally and dy vertically, then the Pythagorean theorem gives the tiny arclength ds:

Here, y' = dy/dx is the derivative/slope of y(x) with respect to x. The total length S of the curve, often referred to as the "action", is found by integrating these arclengths:

Notice that, unlike ordinary functions that take numbers as inputs, S[y] takes an entire curve y(x) as input and outputs a single number, S. Such a function is called a functional.

Our goal is to find the curve y(x) that minimizes the length S[y]. Instead of asking how the function (or "functional") S[y] changes for some change in x, we must ask how the total path length changes if we slightly deform the curve y(x) itself. This is the central idea behind a branch of calculus called "Calculus of Variations".

Digging Deeper

Let's predict that the shortest path between two curves is y(x). Without knowing that y(x) is a straight line (we haven't proved this yet!), our first prediction is probably incorrect. So, let's say the true ideal path is a nearby path y(x) + ε*η(x), where η(x) is the shape of the function that corrects our original prediction, and ε scales this correction function (i.e. controls how large the deformation is). To ensure that the endpoints remain the same, we require η(x1) = η(x2) = 0.

Substituting this new path into the expression for S[y] above, we get the path length S as a function of ε:

Notice that if ε = 0, we recover the original path y(x). If the original path does minimize S, then S has a minimum at ε = 0. That is,

So, to find the minimizing curve y(x), we find a curve y(x) such that dS/dε = 0 at ε = 0. Bringing the derivative inside of the integral for S[y] and using the chain rule, this means:

We can integrate this by parts using the substitutions:

Since we originally defined η(x1) = η(x2) = 0, the first term on the right disappears. We're left with:

Now, this must be true for any η(x) we choose, meaning:

Squaring both sides and rearranging to solve for y', we get that:

Hence, y(x) has a constant slope and is therefore a straight line. Note that there are multiple ways to prove that the shortest path between two points is a straight line (such as using the Triangle Inequality or algebraic arguments under Euclidean geometry), but this method of using least action has proven to have much deeper physical significance to more sophisticated questions, as discussed below.

Important Distinction

As many of you have probably noticed, this argument only holds in flat Euclidean space (the space most of us are used to, like the xyz plane). In curved spaces, the analogue of a straight line is called a geodesic.

The general idea is still the same, but the formula for distance changes depending on the geometry of the space.

For example, consider the surface of the Earth. If you want the shortest route from Toronto to London while remaining on Earth's surface, you don't follow what looks like a straight line on a flat map. You follow part of a "great circle".

On a sphere of radius R, for example, the infinitesimal distance is given using spherical coordinates:

So instead of minimizing the flat-space length

we minimize:

Applying the same calculus-of-variations machinery gives the geodesics of the sphere, which turn out to be great circles.

This becomes especially interesting in general relativity. Spacetime itself is curved, and free-falling objects follow geodesics through that curved spacetime. In that sense, Earth's orbit around the Sun can be thought of not as Earth being forced away from a straight path, but as Earth following the natural geodesic of curved spacetime.

Deeper Physical Significance

What we proved above is a remarkable fact, as the idea of minimizing the "action" functional is a fundamental characteristic of systems in nature. Instead of minimizing distance, physical systems obey the same mathematics to minimize quantities involving energy and time. For example, a hanging rope will form a shape called a catenary (a type of hyperbolic cosine function) to minimize its gravitational potential energy, not a parabola (which approximates the motion of projectiles) as one might think.

Similarly, light bends when moving between materials like air and water to minimize travel time.

So, never doubt the power of seemingly "simple" or intuitive mathematical results, as this elementary question of minimizing distance between two points uses the same mathematics that governs planetary motion, quantum fields, and spacetime itself!

Please Wait...