Calculus of Variations: Variations and the First Variation
CV01 established the setting of the classical variational problem. A functional acts on an admissible
class of functions, and local optimality must be defined relative to a chosen neighborhood in function
space. The next question is the analogue of ordinary differentiation:
If one perturbs an admissible function slightly, how does the scalar value of J change?
The central idea is to convert an infinite-dimensional problem into a one-parameter problem. Instead of
trying to vary a function in all possible ways at once, one chooses a single admissible perturbation shape
η(x) and studies the family
For fixed y and η, the functional becomes an ordinary scalar function of 𝜖,
The derivative of Φ at 𝜖 = 0 is the first variation. It plays the role that f′(x∗) plays in ordinary
calculus. If y is a local extremum, the first variation must vanish for every admissible direction
η.
This entry makes that construction precise, derives the standard first variation formula for first-order
integral functionals, and works several examples in full detail. The next two entries will then use
integration by parts and the Fundamental Lemma to pass from the vanishing of the first variation to the
Euler–Lagrange equation.
1 Learning objectives
After this entry, the reader should be able to
- define a variation and an admissible one-parameter family y𝜖 = y + 𝜖η;
- explain why endpoint conditions on y induce endpoint conditions on η;
- define the scalar reduction Φ(𝜖) = J[y + 𝜖η];
- define the first variation δJ[y; η] as a directional derivative;
- derive the first variation formula for J[y] = ∫
abF(x,y,y′) dx;
- state clearly where differentiation under the integral sign is used;
- prove that the first variation is linear in η;
- compute first variations directly by expansion in 𝜖;
- recognize that stationarity requires δJ[y; η] = 0 for every admissible η;
- distinguish the first variation from the second variation; and
- see how the first variation formula is poised for integration by parts and the Fundamental
Lemma.
2 Variations as admissible perturbation directions
Let 𝒜 be an admissible class for a first-order problem. A variation is a comparison direction η(x) used to
perturb a candidate function y(x).
The classical perturbed family is
where 𝜖 ∈ ℝ is small.
The word “small” has two roles. First, it keeps y𝜖 near y in whatever topology is relevant to the
problem. Second, it allows one to use ordinary differentiation with respect to the scalar parameter
𝜖.
A variation must be chosen so that, for sufficiently small 𝜖, the perturbed curve remains admissible.
Thus the admissibility conditions on y induce admissibility conditions on η.
2.1 Fixed endpoint problems
Suppose the admissible class is
If y𝜖 = y + 𝜖η is to satisfy the same endpoint conditions for all sufficiently small 𝜖, then
Substituting the perturbed family gives
Because y is already admissible, y(a) = A and y(b) = B, so one obtains
Thus, for fixed-endpoint problems, admissible variations vanish at the endpoints.
Figure. A typical one-parameter family y𝜖 = y + 𝜖η for a fixed-endpoint problem. The
comparison curves share the same endpoints because the variation η vanishes at x = a and
x = b.
2.2 Other constraints
If the admissible class includes different constraints, the variation must respect them in the
corresponding linearized sense.
- For periodic constraints y(a) = y(b), one requires η(a) = η(b).
- For an integral constraint such as ∫
abG(x,y,y′) dx = C, not every η is immediately
admissible; this issue leads to constrained variations and Lagrange multipliers in later entries.
- For inequality or obstacle constraints, the admissible directions may be one-sided rather
than arbitrary.
For the present entry, fixed-endpoint unconstrained problems are sufficient to show the core
mechanism.
3 From an infinite-dimensional problem to an ordinary derivative
The first conceptual simplification in the calculus of variations is the map
Once y and η are fixed, the functional no longer acts on an entire class of functions. It acts on a
one-parameter subfamily. Along that subfamily the problem is ordinary calculus.
Figure. The standard reduction. A direction η and scalar amplitude 𝜖 produce a comparison
family y + 𝜖η. The functional then becomes the ordinary scalar function Φ(𝜖) = J[y + 𝜖η].
Differentiating at 𝜖 = 0 yields the first variation.
3.1 Definition of the first variation
Let J be a functional and suppose the derivative exists. The first variation of J at y in the direction η
is
Equivalently, if
then
This is the classical Gâteaux or directional derivative viewpoint. One does not yet require the stronger
uniform approximation property associated with the Fréchet derivative. For deriving the
Euler–Lagrange equation, the directional viewpoint is the natural starting point.
3.2 Stationarity as an infinite-dimensional Fermat principle
In ordinary calculus, if x∗ is an interior local extremum of a differentiable function f, then
In several variables, if x∗ is an interior local extremum of a smooth function f(x), then
The analogous variational principle is:
If y∗ is a local extremum of J, then for every admissible variation η,
The crucial phrase is “for every admissible variation.” A single vanishing value of δJ[y; η] along one
direction says very little. A local extremum must be stationary with respect to all admissible first-order
perturbations.
4 First variation of a first-order integral functional
Consider the classical functional
where F is continuously differentiable in its arguments and y,η are sufficiently smooth for the
expressions below to make sense.
Set
Then
4.1 Main formula
proposition. Assume that F ∈ C1 and that differentiation under the integral sign is justified for the
perturbed family. Then the first variation is
where Fy and Fy′ are evaluated along the unperturbed curve (x,y(x),y′(x)).
4.2 Proof
Differentiate Φ(𝜖) with respect to 𝜖:
Assuming differentiation may pass under the integral sign,
Now apply the chain rule to the integrand. Since x is independent of 𝜖,
Hence
Evaluating at 𝜖 = 0 gives
which is the claimed formula.
4.3 Why the proof is important
The proof shows exactly where the machinery comes from. The first variation is not a mysterious formal
symbol. It is an ordinary derivative with respect to 𝜖 after one restricts the functional to a
one-parameter family.
The proof also isolates the analytic step that needs hypotheses: differentiation under the integral sign. In
classical coursework one usually assumes enough smoothness for this to be valid. More advanced
functional analysis develops sharper conditions, but the basic variational idea is already visible in the
classical setting.
5 Linearity of the first variation
For fixed y, the map η
δJ[y; η] is linear.
Proposition. Let α,β ∈ ℝ and let η1,η2 be admissible variations. Then
Proof. Apply the first variation formula:
Distribute and use linearity of the integral:
Therefore
This linearity is the reason one can think of δJ[y; ⋅] as an analogue of a differential or covector acting on
perturbation directions.
6 Direct computation by 𝜖-expansion
Although the general formula is fundamental, students should be able to compute first variations
directly from the definition. Expanding J[y + 𝜖η] in powers of 𝜖 is often the clearest way to see what is
happening.
6.1 Example 1: the standard energy functional around the straight line
Consider
with fixed endpoints
Take the candidate
and the admissible variation
which satisfies η(0) = η(1) = 0.
The perturbed family is
so
Substitute into J:
Expand the square:
The middle integral vanishes and the last integral equals 1∕2, hence
Therefore
This shows that the straight line is stationary in the direction η(x) = sin(πx). In fact it is stationary for
every admissible η, which will become completely transparent after integration by parts in
CV04.
Figure. For the family y𝜖 = x + 𝜖 sin(πx), the scalarized functional is Φ(𝜖) = 1 + (π2∕2)𝜖2. The
tangent at 𝜖 = 0 is horizontal, so the first variation vanishes.
6.2 Example 2: use the general formula
Now compute the same first variation from the formula. Here
Therefore
The first variation at a general admissible curve y is
For y(x) = x, one has y′ = 1, so
Thus the straight line is stationary for every fixed-endpoint variation, as expected.
6.3 Example 3: a nonstationary admissible curve
Still with
consider the admissible curve
which satisfies y(0) = 0 and y(1) = 1.
Choose the admissible variation
Then
Using the formula,
Compute the integral:
Because the first variation is not zero, the curve y = x2 is not stationary. So one does not need the full
Euler–Lagrange equation merely to rule out a candidate. A single admissible direction with nonzero first
variation is already enough.
6.4 Example 4: a functional containing both y and y′
Consider
with fixed endpoints y(0) = y(1) = 0 and candidate y(x) = 0.
Let η be any admissible variation, so η(0) = η(1) = 0. Then
Hence
so
This example is useful because it shows that the first variation alone detects stationarity,
not the full classification. Here the absence of a linear term already suggests that y = 0
should be minimizing, but one needs second-order information to establish that rigorously in
general.
7 The bridge to Euler–Lagrange
The first variation formula still contains η′:
The next step is to remove the derivative from the variation by integration by parts:
Substituting gives
For fixed-endpoint variations, η(a) = η(b) = 0, so the boundary term drops out and one
obtains
Now the structure is clear. If y is stationary, then this integral vanishes for every admissible variation η.
The Fundamental Lemma will then imply
which is the Euler–Lagrange equation.
This logical order matters:
CV03 is devoted to the Fundamental Lemma because that final implication is a nontrivial theorem, not
a formal trick.
8 Common misconceptions
8.1 Misconception 1: the variation η is itself the new curve
No. The perturbed curve is y + 𝜖η. The function η provides the shape of the perturbation, while 𝜖
controls its size.
8.2 Misconception 2: if one example gives δJ = 0, the curve is stationary
No. Stationarity requires
A single direction is not enough.
8.3 Misconception 3: the first variation already proves a minimum
Not in general. Vanishing first variation is a necessary condition for an interior local extremum, but not
a sufficient one. Classification belongs to second-variation theory.
8.4 Misconception 4: the first variation formula is obtained by formally “varying” symbols
The formula does have a compact formal notation, but its basis is ordinary calculus on the scalar
function Φ(𝜖) = J[y + 𝜖η].
8.5 Misconception 5: differentiation under the integral sign is automatic
One needs hypotheses to justify it. In introductory classical treatments the required regularity is usually
assumed from the start.
9 Compact definition table
|
|
| Object | Meaning |
|
|
| η(x) | variation or perturbation direction |
|
|
| y𝜖 = y + 𝜖η | one-parameter family of nearby curves |
|
|
| Φ(𝜖) = J[y + 𝜖η] | scalarized functional along one family |
|
|
| δJ[y; η] | first variation, equal to Φ′(0) |
|
|
| δ2J | second variation, based on Φ′′(0) |
|
|
| stationary function | one for which δJ[y; η] = 0 for all admissible η |
|
|
10 What CV03 and CV04 add
CV03 will prove the Fundamental Lemma of the Calculus of Variations. That lemma tells us that
if
for every smooth test function η with compact support or vanishing endpoints, then g(x) = 0
identically.
CV04 will combine that lemma with the integration-by-parts form of the first variation to obtain the
Euler–Lagrange equation in its standard classical form.
Thus CV02 provides the derivative computation, CV03 provides the theorem that converts an integral
identity into a pointwise statement, and CV04 assembles those pieces into the central necessary
condition of the subject.
11 Summary
The first variation is the directional derivative of a functional. Starting from a candidate y and an
admissible variation η, one builds the family y + 𝜖η and defines
For the classical first-order functional
the first variation is
This quantity is linear in the variation direction and must vanish for every admissible variation if y is a
stationary function.
The first variation therefore plays the same conceptual role in the calculus of variations that the ordinary
derivative plays in finite-dimensional calculus. Its integration-by-parts form is the gateway to the
Fundamental Lemma and the Euler–Lagrange equation.
12 References and further reading
The following references are standard classical entry points into the subject. They are listed here for
mathematical orientation; the present article remains self-contained.
- I. M. Gelfand and S. V. Fomin, Calculus of Variations.
- B. van Brunt, The Calculus of Variations.
- C. Fox, An Introduction to the Calculus of Variations.
- L. C. Evans, Partial Differential Equations, for the modern weak and functional-analytic
viewpoint.
- H. Goldstein, C. Poole, and J. Safko, Classical Mechanics, for the action principle
connection.