Lagrange multipliers on manifolds


We discuss in this article the theoretical aspects of the Lagrange multiplier method.

To enhance understanding, proofs and intuitive explanations of the Lagrange multipler method will be given from several different viewpoints, both elementary and advanced.

1 Statements of theorem

Let N be a n-dimensional differentiable manifold (without boundary), and f:N→ℝ, and gi:N→ℝ, for i=1,…,k, be continuously differentiable. Set M=⋂i=1kgi-1⁢({0}).

1.1 Formulation with differential forms

Theorem 1.

Suppose d⁢gi are linearly independentMathworldPlanetmath at each point of M. If p∈M is a local minimumMathworldPlanetmath or maximum point of f restricted to M, then there exist Lagrange multipliers λ1,…,λk∈R, depending on p, such that

d⁢f⁢(p)=λ1⁢d⁢g1⁢(p)+⋯+λk⁢d⁢gk⁢(p).

Here, d denotes the exterior derivativeMathworldPlanetmath.

Of course, as in one-dimensional calculus, the condition d⁢f⁢(p)=∑iλi⁢d⁢gi⁢(p) by itself does not guarantee p is a minimum or maximum point, even locally.

1.2 Formulation with gradients

The version of Lagrange multipliers typically used in calculus is the special case N=ℝn in Theorem 1. In this case, the conclusionMathworldPlanetmath of the theorem can also be written in terms of gradientsMathworldPlanetmath instead of differential forms:

Theorem 2.

Suppose ∇⁡gi are linearly independent at each point of M. If p∈M is a local minimum or maximum point of f restricted to M, then there exist Lagrange multipliers λ1,…,λk∈R, depending on p, such that

∇⁡f⁢(p)=λ1⁢∇⁡g1⁢(p)+⋯+λk⁢∇⁡gk⁢(p).

This formulation and the first one are equivalentMathworldPlanetmathPlanetmathPlanetmathPlanetmathPlanetmath since the 1-form d⁢f can be identified with the gradient ∇⁡f, via the formulaMathworldPlanetmathPlanetmath ∇⁡f⁢(p)⋅v=d⁢f⁢(p;v)=d⁢fp⁢(v).

1.3 Formulation with tangent maps

The functions gi can also be coalesced into a vector-valued functionPlanetmathPlanetmath g:N→ℝk. Then we have:

Theorem 3.

Let g=(g1,…,gk):N→Rk. Suppose the tangent map D⁡g is surjectivePlanetmathPlanetmath at each point of M. If p∈M is a local minimum or maximum point of f restricted to M, then there exists a Lagrange multiplier vector λ∈(Rk)*, depending on p, such that

D⁡f⁢(p)=D⁡g⁢(p)*⁢λ.

Here, D⁡g⁢(p):(Rk)*→(Tp⁢N)* denotes the pullback of the linear transformation (http://planetmath.org/DualHomomorphism) D⁡g⁢(p):Tp⁢N→Rk.

If D⁡g is represented by its Jacobian matrix, then the condition that it be surjective is equivalent to its Jacobian matrix having full rank.

Note the deliberate use of the space (ℝk)* instead of ℝk — to which the former is isomorphic to — for the Lagrange multiplier vector. It turns out that the Lagrange multiplier vector naturally lives in the dual spaceMathworldPlanetmathPlanetmathPlanetmath and not the original vector spaceMathworldPlanetmath ℝk. This distinction is particularly important in the infinite-dimensional generalizationsPlanetmathPlanetmath of Lagrange multipliers. But even in the finite-dimensional setting, we do see hints that the dual space has to be involved, because a transposeMathworldPlanetmath is involved in the matrix expression for Lagrange multipliers.

If the expression D⁡g⁢(p)*⁢λ is written out in coordinatesMathworldPlanetmathPlanetmath, then it is apparent that the components λi of the vector λ are exactly those Lagrange multipliers from Theorems 1 and 2.

2 Proofs

The proof of the Lagrange multiplier theorem is surprisingly short and elegant, when properly phrased in the languagePlanetmathPlanetmath of abstract manifolds and differential forms.

However, for the benefit of the readers not versed in these topics, we provide, in addition to the abstract proof, a concrete translationMathworldPlanetmathPlanetmath of the argumentsMathworldPlanetmath in the more familiar setting N=ℝn.

2.1 Beautiful abstract proof

Proof.

Since d⁢gi are linearly independent at each point of M=⋂i=1kgi-1⁢({0}), M is an embedded submanifold of N, of dimensionMathworldPlanetmathPlanetmath m=n-k. Let α:U→M, with U open in ℝm, be a coordinate chart for M such that α⁢(0)=p. Then α*⁢f has a local minimum or maximum at 0, and therefore 0=d⁢(α*⁢f)=α*⁢d⁢f at 0. But α* at p is an isomorphismPlanetmathPlanetmathPlanetmath (Tp⁢M)*→(T0⁢ℝm)*, so the preceding equation says that d⁢f vanishes on Tp⁢M.

Now, by the definition of gi, we have α*⁢gi=0, so 0=d⁢(α*⁢gi)=α*⁢d⁢gi. So like d⁢f, d⁢gi vanishes on Tp⁢M.

In other words, d⁢gi⁢(p) is in the annihilatorMathworldPlanetmathPlanetmathPlanetmath (http://planetmath.org/AnnihilatorOfVectorSubspace) (Tp⁢M)0 of the subspaceMathworldPlanetmathPlanetmathPlanetmath Tp⁢M⊆Tp⁢N. Since Tp⁢M has dimension m=n-k, and Tp⁢N has dimension n, the annihilator (Tp⁢M)0 has dimension k. Now d⁢gi⁢(p)∈(Tp⁢M)0 are linearly independent, so they must in fact be a basis for (Tp⁢M)0. But we had argued that d⁢f⁢(p)∈(Tp⁢M)0. Therefore d⁢f⁢(p) may be written as a unique linear combinationMathworldPlanetmath of the d⁢gi⁢(p):

d⁢f⁢(p)=λ1⁢d⁢g1⁢(p)+⋯+λk⁢d⁢gk⁢(p).∎

The last paragraph of the previous proof can also be rephrased, based on the same underlying ideas, to make evident the fact that the Lagrange multiplier vector lives in the dual space (ℝk)*.

Alternative argument..

A general theorem in linear algebra states that for any linear transformation L, the image of the pullback L* is the annihilator of the kernel of L. Since ker⁡D⁡g⁢(p)=Tp⁢M and d⁢f⁢(p)∈(Tp⁢M)0, it immediately follows that λ∈(ℝk)* exists such that d⁢f⁢(p)=D⁡g⁢(p)*⁢λ. ∎

Yet another proof could be devised by observing that the result is obvious if N=ℝn and the constraint functions are just coordinate projections on ℝn:

gi⁢(y1,…,yn)=yi,i=1,…,k.

We clearly must have ∂⁡f/∂⁡yk+1=⋯=∂⁡f/∂⁡yn=0 at a point p that minimizes f⁢(y) over y1=⋯=yk=0. The general case can be deduced to this by a coordinate change:

Alternate argument..

Since d⁢gi are linearly independent, we can find a coordinate chart for N about the point p, with coordinate functions y1,…,yn:N→ℝ such that yi=gi for i=1,…,k. Then

d⁢f =∂⁡f∂⁡y1⁢d⁢y1+⋯+∂⁡f∂⁡yn⁢d⁢yn
=∂⁡f∂⁡g1⁢d⁢g1+⋯+∂⁡f∂⁡gk⁢d⁢gk+∂⁡f∂⁡yk+1⁢d⁢yk+1+⋯+∂⁡f∂⁡yn⁢d⁢yn,

but ∂⁡f/∂⁡yk+1=⋯=∂⁡f/∂⁡yn=0 at the point p. Set λi=∂⁡f/∂⁡gi at p. ∎

2.2 Clumsy, but down-to-earth proof

Proof.

We assume that N=ℝn. Consider the list vector g=(g1,…,gk) discussed earlier, and its Jacobian matrix D⁡g in EuclideanPlanetmathPlanetmath coordinates. The ith row of this matrix is

[∂⁡gi∂⁡x1…∂⁡gi∂⁡xn]=(∇⁡gi)T.

So the matrix D⁡g has full rank (i.e. rank⁡D⁡g=k) if and only if the k gradients ∇⁡gi are linearly independent.

Consider each solution q∈M of g⁢(q)=0. Since D⁡g has full rank, we can apply the implicit function theorem, which states that there exist smooth solution parameterizations α:U→M around each point q∈M. (U is an open set in ℝm, m=n-k.) These α are the coordinate charts which give to M=g-1⁢({0}) a manifold structureMathworldPlanetmath.

We now consider specially the point q=p; without loss of generality, assume α⁢(0)=p. Then f∘α is a function on Euclidean spaceMathworldPlanetmath having a local minimum or maximum at 0, so its derivativePlanetmathPlanetmath vanishes at 0. Calculating by the chain ruleMathworldPlanetmath, we have 0=D⁡(f∘α)⁢(0)=D⁡f⁢(p)⋅D⁡α⁢(0). In other words, ker⁡D⁡f⁢(p)⊇range of ⁢D⁡α⁢(0)=Tp⁢M. Intuitively, this says that the directional derivativesMathworldPlanetmath at p of f lying in the tangent spaceMathworldPlanetmathPlanetmath Tp⁢M of the manifold M vanish.

By the definition of g and α, we have g∘α=0. By the chain rule again, we derive 0=D⁡g⁢(p)⋅D⁡α⁢(0).

Let the columns of D⁡α⁢(0) be the column vectors v1,…,vm, which span the m-dimensional space Tp⁢M, and look at the matrix equation 0=D⁡f⁢(p)⋅D⁡α⁢(0) again. The equation for each entry of this matrix, which consists of only one row, is:

∇⁡f⁢(p)⋅vj=0,j=1,…,m.

In other words, ∇⁡f⁢(p) is orthogonalMathworldPlanetmathPlanetmathPlanetmath to v1,…,vm, and hence it is orthogonal to the entire tangent space Tp⁢M.

Similarly, the matrix equation 0=D⁡g⁢(p)⋅D⁡α⁢(0) can be split into individual scalar equations:

∇⁡gi⁢(p)⋅vj=0,i=1,…,k,j=1,…,m.

Thus ∇⁡gi⁢(p) is orthogonal to Tp⁢M. But ∇⁡gi⁢(p) are, by hypothesisMathworldPlanetmath, linearly independent, and there are k of these gradients, so they must form a basis for the orthogonal complementMathworldPlanetmath of Tp⁢M, of n-m=k dimensions. Hence ∇⁡f⁢(p) can be written as a unique linear combination of ∇⁡gi⁢(p):

∇⁡f⁢(p)=λ1⁢∇⁡g1⁢(p)+⋯+λk⁢∇⁡gk⁢(p).∎

3 Intuitive interpretations

We now discuss the intuitive and geometric interpretationsMathworldPlanetmath of Lagrange multipliers.

3.1 Normals to tangent hyperplanes

Each equation gi=0 defines a hypersurface Mi in ℝn, a manifold of dimension n-1. If we consider the tangentPlanetmathPlanetmathPlanetmath hyperplaneMathworldPlanetmathPlanetmath at p of these hypersurfaces, Tp⁢Mi, the gradient ∇⁡gi⁢(p) gives the normal vectorMathworldPlanetmath to these hyperplanes.

The manifold M is the intersectionMathworldPlanetmath of the hypersurfaces Mi. Presumably, the tangent space Tp⁢M is the intersection of the Tp⁢Mi, and the subspace perpendicularMathworldPlanetmathPlanetmath to Tp⁢M would be spanned by the normals ∇⁡gi⁢(p). Now, the direction derivatives at p of f with respect to each vector in Tp⁢M, as we have proved, vanish. So the direction of ∇⁡f⁢(p), the direction of the greatest change in f at p, should be perpendicular to Tp⁢M. Hence ∇⁡f⁢(p) can be written as a linear combination of the ∇⁡gi⁢(p).

Note, however, that this geometric picture, and the manipulations with the gradients ∇⁡f⁢(p) and ∇⁡gi⁢(p), do not carry over to abstract manifolds. The notions of gradients and normals to surfaces depend on the inner productMathworldPlanetmath structure of ℝn, which is not present in an abstract manifold (without a Riemannian metricMathworldPlanetmath).

On the other hand, this explains the mysterious appearance of annihilators in the last paragraph of the abstract proof. Annihilators and dual space theory serve as the proper tools to formalize the manipulations we made with the matrix equations 0=D⁡f⁢(p)⋅D⁡α⁢(0) and 0=D⁡g⁢(p)⋅D⁡α⁢(0), without resorting to Euclidean coordinates, which, of course, are not even defined on an abstract manifold.

3.2 With infinitesimals

If we are willing to interpret the quantities d⁢f and d⁢gi as infinitesimalsMathworldPlanetmathPlanetmath, even the abstract version of the result has an intuitive explanation. Suppose we are at the point p of the manifold M, and consider an infinitesimal movement Δ⁢p about this point. The infinitesimal movement Δ⁢p is a vector in the tangent space Tp⁢M, because, near p, M looks like the linear space Tp⁢M. And as p moves, the function f changes by a corresponding infinitesimal amount d⁢f that is approximately linear in Δ⁢p.

Furthermore, the change d⁢f may be decomposed as the sum of a change as p moves along the manifold M, and a change as p moves out of the manifold M. But if f has a local minimum at p, then there cannot be any change of f along M; thus f only changes when moving out of M. Now M is described by the equations gi=0, so a movement out of M is described by the infinitesimal changes d⁢gi. As d⁢f is linear in the change Δ⁢p, we ought to be able to write it as a weighted sum of the changes d⁢gi. The weights are, of course, the Lagrange multipliers λi.

The linear algebra performed in the abstract proof can be regarded as the precise, rigorous translation of the preceding argument.

3.3 As rates of substitution

Observe that the formula for Lagrange multipliers is formally very similarMathworldPlanetmathPlanetmath to the standard formula for expressing a differential form in terms of a basis:

d⁢f⁢(p)=∂⁡f∂⁡y1⁢d⁢y1+⋯+∂⁡f∂⁡yk⁢d⁢yk.

In fact, if d⁢gi⁢(p) are linearly independent, then they do form a basis for (Tp⁢M)0, that can be extended to a basis for (Tp⁢N)*. By the uniqueness of the basis representation, we must have

λi=∂⁡f∂⁡gi.

That is, λi is the differentialMathworldPlanetmath of f with respect to changes in gi.

In applications of Lagrange multipliers to economic problems, the multipliers λi are rates of substitution — they give the rate of improvement in the objective function f as the constraints gi are relaxed.

4 Stationary points

In applications, sometimes we are interested in finding stationary points p of f — defined as points p such that d⁢f vanishes on Tp⁢M, or equivalently, that the Taylor expansion of f at p, under any system of coordinates for M, has no terms of first order. Then the Lagrange multiplier method works for this situation too.

The following theorem incorporates the more general notion of stationary points.

Theorem 4.

Let N be a n-dimensional differentiable manifold (without boundary), and f:N→R, gi:N→R, for i=1,…,k, be continuously differentiable. Suppose p∈M=⋂i=1kgi-1⁢({0}), and d⁢gi⁢(p) are linearly independent.

Then p is a stationary point (e.g. a local extremum point) of f restricted to M, if and only if there exist λ1,…,λk∈R such that

d⁢f⁢(p)=λ1⁢d⁢g1⁢(p)+⋯+λk⁢d⁢gk⁢(p).

The Lagrange multipliers λi, which depend on p, are unique when they exist.

In this formulation, M is not necessarily a manifold, but it is one when intersected with a sufficiently small neighborhood about p. So it makes sense to talk about Tp⁢M, although we are abusing notation here. The subspace in question can be more accurately described as the annihilated subspace of span⁡{d⁢gi⁢(p)}.

It is also enough that d⁢gi be linearly independent only at the point p. For d⁢gi are continuousMathworldPlanetmathPlanetmath, so they will be linearly independent for points near p anyway, and we may restrict our viewpoint to a sufficiently small neighborhood around p, and the proofs carry through.

The proof involves only simple modifications to that of Theorem 1 — for instance, the converse implication follows because we have already proved that the d⁢gi⁢(p) form a basis for the annihilator of Tp⁢M, independently of whether or not p is a stationary point of f on M.

References

  • 1 Friedberg, Insel, Spence. Linear Algebra. Prentice-Hall, 1997.
  • 2 David Luenberger. Optimization by Vector Space Methods. John Wiley & Sons, 1969.
  • 3 James R. Munkres. Analysis on Manifolds. Westview Press, 1991.
  • 4 R. Tyrrell Rockafellar. “Lagrange Multipliers and Optimality”. SIAM Review. Vol. 35, No. 2, June 1993.
  • 5 Michael Spivak. Calculus on Manifolds. Perseus Books, 1998.
Title Lagrange multipliers on manifolds
Canonical name LagrangeMultipliersOnManifolds
Date of creation 2013-03-22 15:25:45
Last modified on 2013-03-22 15:25:45
Owner stevecheng (10074)
Last modified by stevecheng (10074)
Numerical id 24
Author stevecheng (10074)
Entry type Topic
Classification msc 58C05
Classification msc 49-00
Related topic Manifold