In the previous part we have constructed our two fundamental solutions $y_1$ and $y_2$ to the ODE $$ -y_j^{\prime\prime} + q(x)y_j = \lambda y_j, $$ with initial data $$ y_1(0) = y_2^\prime(0) = 1,\quad y_1^\prime(0) = y_2(0) = 0. $$
The goal of this section is to establish useful facts about the $y_j$, for we need to use them in the later sections to establish meaningful results. These facts come in the form of estimates, knowledge about the (partial) derivatives of the $y_j$, as well as analyticity properties.
Understandably, this is rather dry, and reads like a collection of rather disparate results. One can probably skip this section and refer back to it as needed.
But first, recall that we mentioned that for our ODE with general initial data, we can write the solution in terms of $y_1$ and $y_2$. We will show this now. This also serves as a good introduction to the following useful object:
Definition 2. The Wronskian of two differentiable functions $f$ and $g$ is the function $$ [f,g] := \begin{vmatrix}f && g \\ f^\prime &&g^\prime\end{vmatrix} = fg^\prime - f^\prime g. $$
Along with the Wronskian, it will be necesary to introduce the following Wronskian identity.
Lemma 2. One has $[y_1, y_2] = 1$.
Proof. Just compute $$ [y_1, y_2]^\prime = y_1y_2^{\prime\prime} - y_1^{\prime\prime}y_2 = y_1(q-\lambda)y_2 - (q-\lambda)y_1y_2 = 0 \text{ a.e.}, $$ whence by continuity $$ [y_1, y_2](x) = [y_1, y_2](0) = 1. \quad \square $$
With that, we have the following theorem, which is essentially a generalisation of Lemma 1 in the previous part.
Theorem 2. Let $f\in L^2(\mathbb{C})$ and $a,b\in\mathbb{C}$. Then there exists a unique solution to the inhomogenuous ODE $$-y^{\prime\prime} + q(x)y = \lambda y - f(x), \quad 0\leq x\leq 1,$$ satisfying $y(0) = a$, $y^\prime(0)=b$. It is given by the explicit formula $$ y(x) = ay_1(x) + by_2(x) + \int_0^x (y_1(t)y_2(x) - y_1(x)y_2(t))f(t)\,dt. $$
The proof is literally the same as Lemma 1, except that we need to use the Wronskian identity. Therefore, we will omit it. As a corollary, when $f(x)=0$ (which is the case we really care about), the unique solution is $y(x) = y(0)y_1(x) + y^\prime(0)y_2(x)$. In fact, if this solution has a double root (i.e. $y(a) = y^\prime(a) = 0$ for some $a\in[0,1]$), the solution must vanish identically on $[0,1]$, for $$ 0 = \begin{pmatrix}y(a) \\ y^\prime(a) \end{pmatrix} = \begin{pmatrix}y_1(x) && y_2(a) \\ y_1^\prime(a) && y_2^\prime(a) \end{pmatrix} \begin{pmatrix}y(0) \\ y^\prime(0) \end{pmatrix}, $$ and the matrix is invertible, so $y(0) = y^{\prime}(0)= 0$, whence $y=0$. $\square$
Recall that when $q=0$, we have $$ \begin{align*} y_1(x) = \cos(\sqrt{\lambda}x) = \sum_{n\geq 0}\frac{(-1)^n}{(2n)!}x^{2n}\lambda^n, \\y_2(x) = \frac{\sin(\sqrt{\lambda}x)}{\sqrt{\lambda}} = \sum_{n\geq 0}\frac{(-1)^n}{(2n+1)!}x^{2n+1}\lambda^n. \end{align*} $$ These functions are entire on $\lambda$, which is evident once we expand out the Taylor series of $y_1$ and $y_2$ as above. We can generalise this to nonzero $q$.
Theorem 3. (Analyticity Properties)
- For each $x\in [0,1]$ the maps $y_1(x, \lambda, q)$ and $y_2(x, \lambda, q)$ are analytic on $\mathbb{C}\times L^2_\mathbb{C}$. Furthermore, $y_1(x, \lambda, q)$ and $y_2(x, \lambda, q)$ are real-valued on $\mathbb{R}\times L^2_\mathbb{R}$.
- Each $y_j(\cdot, \lambda, q)$ is analytic as a map from $\mathbb{C}\times L^2_\mathbb{C}$ into $H^2_\mathbb{C}$.
Recall that $H^2_\mathbb{C}$ is the Sobolev space of complex-valued $L^2$ functions, whose derivative and second derivative exist in $L^2$.
We will just prove this for $y_1$, because the proof for $y_2$ is exactly the same.
Proof.
- It suffices to show that $y_1$ is continuously differentiable, and then invoke the Cauchy integral formula. Furthermore, by local uniform convergence, it suffices to check that each $C_n$ is continuously differentiable. But $$ C_n(x, \lambda, q) = \int_{0\leq t_1\leq\dots\leq t_{n+1}=x}\underbrace{c_\lambda(t_1)\prod_{i=1}^{n}s_\lambda(t_{i+1}-t_i)}_{\text{$C^1$ in $\lambda$}}\prod_{i=1}^n q(t_i)\,dt_1\dots\,dt_n, $$ and $C_n(x, \lambda, q) = C_n(x, \lambda, q_1, \dots, q_n)|_{q_1=\dots=q_n=q}$ is bounded and multilinear in the $q_i$, and so it is also $C^1$ in $q$. This prove analyticity. Now if $\lambda$ is real and $q$ is real-valued, we can take the complex-conjugate of the integral equation $$ y_1(x, \lambda, q) = c_\lambda(x) + \int_0^x s_\lambda(x-t)q(t)y_1(x, \lambda, q)\,dt, $$ to obtain $$ \overline{y_1(x, \lambda, q)} = c_\lambda(x) + \int_0^x s_\lambda(x-t)q(t)\overline{y_1(x, \lambda, q)}\,dt. $$ Using the same Grönwall-style argument as in Theorem 1, we will obtain $y_1 = \overline{y_1}$. $\square$
- It is evident that each $C_n(\cdot, \lambda, q)$ is anlaytic from $\mathbb{C}\times L^2_\mathbb{C}$ into $C^0$. By local uniform convergence this carries to $y_1$, and by the integral equation this carries to $y_1^\prime$. The interval $[0,1]$ being bounded, the supremum norm is stronger than the $L^2$-norm, and so $y_1$ and $y_1^\prime$ are in fact analytic into $L^2_\mathbb{C}([0,1])$. Finally, since $y^{\prime\prime}=(q-\lambda)y$, we also have that $y_1^{\prime\prime}$ maps analytically into $L^2_\mathbb{C}$. $\square$
In Theorem 1 we established the bound $$ |y_1(x, \lambda, q)|, |y_2(x, \lambda, q)| \leq \exp(|\operatorname{Im}\sqrt{\lambda}|x + \|q\|\sqrt{x}), $$ where $|q|$ denotes the $L^2$-norm of $q$. Unfortunately, these bounds are not strong enough to use as-is. We will improve on these to get two collections of estimates: one collection of “basic estimates”, which is simple and will be useful for the most part; and an expansion of the $y_j$, which we will use in more desperate times.
Theorem 4. (Basic Estimates) On $[0,1]\times\mathbb{C}\times L^2_\mathbb{C}$ one has $$ \begin{align*} \left|y_1(x, \lambda, q) - \cos(\sqrt{\lambda}x)\right| \leq\frac{1}{|\sqrt{\lambda}|}\exp\left(|\operatorname{Im}\sqrt{\lambda}|x + \|q\|\sqrt{x}\right),\\ \left|y_2(x, \lambda, q) - \frac{\sin(\sqrt{\lambda}x)}{\sqrt{\lambda}}\right| \leq\frac{1}{|\lambda|}\exp\left(|\operatorname{Im}\sqrt{\lambda}|x + \|q\|\sqrt{x}\right),\\ \end{align*} $$ and $$ \begin{align*} \left|y_1^\prime(x, \lambda, q) + \sqrt{\lambda}\sin(\sqrt{\lambda}x)\right| \leq \|q\|\exp\left(|\operatorname{Im}\sqrt{\lambda}|x + |q|\sqrt{x}\right),\\ \left|y_2^\prime(x, \lambda, q) - \cos(\sqrt{\lambda}x)\right| \leq\frac{\|q\|}{|\sqrt{\lambda}|}\exp\left(|\operatorname{Im}\sqrt{\lambda}|x + |q|\sqrt{x}\right),\\ \end{align*} $$
Proof. We use the bound $$ |s_\lambda(t)| = \frac{2}{|\sqrt{\lambda}|} \left(\exp(i\sqrt{\lambda}t) - \exp(-i\sqrt{\lambda}t)\right) = \frac{1}{|\sqrt{\lambda}|}\exp(|\operatorname{Im}\sqrt{\lambda}|t). $$ Now, repeating the proof of Theorem 1 with this bound on one of the $s_\lambda$ gives, for $n\geq 1$, $$ |C_n(x, \lambda, q)| \leq \frac{1}{|\sqrt{\lambda}|}\exp(|\operatorname{Im}\sqrt{\lambda}|x)\frac{\|q\|^n\sqrt{x}^n}{n!}. $$ Therefore summing over $n\geq 1$ yields $$ \begin{align*} \left|y_1(x, \lambda, q) - \cos(\sqrt{\lambda}x)\right| &\leq \frac{1}{|\sqrt{\lambda}|}\exp(|\operatorname{Im}\sqrt{\lambda}|x)\sum_{n\geq 1}\frac{\|q\|^n\sqrt{x}^n}{n!} \\&\leq \frac{1}{|\sqrt{\lambda}|}\exp(|\operatorname{Im}\sqrt{\lambda}|x + \|q\|\sqrt{x}). \end{align*} $$ The bound for $y_2$ is established analogously. Now, differentiating the integral equation for $y_1$ yields $$ \begin{align*} y_1^\prime(x, \lambda, q) &= -\sqrt{\lambda}\sin(\sqrt{\lambda}x) + \frac{d}{dx}\int_0^x \frac{\sin(\sqrt{\lambda}(x-t))}{\sqrt{\lambda}}q(t)y_1(t)\,dt \\&=-\sqrt{\lambda}\sin(\sqrt{\lambda}x) + \int_0^x\cos(\sqrt{\lambda}(x-t)q(t)y_1(t)\,dt. \end{align*} $$ The derivative of the integral can be evaluated by, for instance, the Leibniz integral rule. This gives $$ \begin{align*} &\left|y_1^\prime(x, \lambda, q) + \sqrt{\lambda}\sin(\sqrt{\lambda}x)\right| \\&\leq \int_0^x \exp(|\operatorname{Im}\sqrt{\lambda}|(x-t))|q(t)|\exp(|\operatorname{Im}\sqrt{\lambda}|t + \|q\|\sqrt{t})\,dt \\&= \exp(|\operatorname{Im}\sqrt{\lambda}|x)\int_0^x |q(t)|\exp(\|q\|\sqrt{t})\,dt \\&\leq \|q\|\exp(|\operatorname{Im}\sqrt{\lambda}|x + \|q\|\sqrt{x}). \end{align*} $$ A similar analysis can be done for $y_2^\prime$. $\square$
By applying our new bound for $|s_\lambda|$ on $n+1$ instances of $s_\lambda$ in $C_k$, $k\geq n$, instead of just one, we in fact have $$ y_1(x, \lambda, q) = \cos(\sqrt{\lambda}x) + \sum_{n=1}^NC_n(x, \lambda, q) + \mathcal{O}\left(\frac{\exp(|\operatorname{Im}\sqrt{\lambda|})}{|\sqrt{\lambda}|^{N+1}}\right). $$ Unfortunately, the $C_n$ are rather complicated, and so this bound offers no practical use. We can try to use high-school trigonometry to simplify this expression, but this can only take us so far. However, if we assume that $q$ has some regularity, we can resort to integration by parts. The process is exceedingly tedious, unenlightening, and downright horrible, and we will spare the reader the derivation (and spare the author from the typing), but you can obtain the following estimate if you try hard enough.
Theorem 5. For $q\in H^2_\mathbb{C}$, we have $$ \begin{align*} y_1(x, \lambda, q) &= \cos(\sqrt{\lambda}x) + \frac{\sin(\sqrt{\lambda}x)}{2\sqrt{\lambda}}Q(x) \\&- \frac{\cos(\sqrt{\lambda}x)}{4\lambda}\left(q(x) - q(0) - \frac{1}{2}Q(x)^2\right) + \mathcal{O}\left(\frac{\exp(|\operatorname{Im}\sqrt{\lambda}|x)}{|\sqrt{\lambda}|^3}\right), \end{align*} $$ and $$ \begin{align*} y_2(x, \lambda, q) &= \frac{\sin(\sqrt{\lambda}x)}{\sqrt{\lambda}} - \frac{\cos(\sqrt{\lambda}x)}{2\lambda}Q(x) \\&+ \frac{\sin(\sqrt{\lambda}x)}{4\sqrt{\lambda}^3}\left(q(x) + q(0) - \frac{1}{2}Q(x)^2\right) + \mathcal{O}\left(\frac{\exp(|\operatorname{Im}\sqrt{\lambda}|x)}{|\lambda|^2}\right), \end{align*} $$ where $Q(x) = \int_0^x q(t)\,dt$.
We remark that in the proof of this theorem, to establish the boundedness of the integrals involved, one will use the fact that any $q\in H^2_\mathbb{C}$ is also $C^1$ with absolutely continuous derivative. This is an example of a Sobolev embedding theorem, but we will not pursue it further.
It is a remarkable fact that continuous functions on a closed, bounded subset of Euclidean space attain their maximum and minimum. Unfortunately, $L^2$ is an infinite-dimensional Banach space, and it is not true that continuous maps on a closed, bounded subset of a Banach space attain their maximum and minimum. The drop-in replacement for continuous maps is the notion of a compact map: instead of asking that strongly convergent sequences get mapped to strongly convergent sequences, we ask that weakly convergent sequences get mapped to strongly convergent sequences. This is a much stronger condition, and it turns out, remarkably, that our $y_j$ does satisfy this condition.
Theorem 6. If $q_m\rightharpoonup q$ in $L^2_\mathbb{C}$, then $y_j(x, \lambda, q_m)\to y_j(x, \lambda, q)$ uniformly on bounded subsets of $[0,1]\times \mathbb{C}$. In other words, the $y_j$ are uniformly compact on bounded subsets of $[0,1]\times \mathbb{C}$.
Proof. Since $q_m\rightharpoonup q$, the Banach-Steinhaus theorem gives us a bound $$ \|q\| \leq \sup_m \|q_m\| \leq M < \infty. $$ For any bounded subset $A$ of $[0,1]\times\mathbb{C}$, and $p=q_m$ or $q$, we have $$ |C_n(x, \lambda, q) \leq \frac{1}{n!}\exp(|\operatorname{Im\sqrt{\lambda}|}x)(\|p\|\sqrt{x})^n \leq c\frac{M^n}{n!}. $$ uniformly on $A$. Thus $$ \begin{align*} &|y_1(x, \lambda, q_m) - y_1(x, \lambda, q)| \\&\leq \sum_{n=1}^N |C_n(x, \lambda, q_m) - C_n(x, \lambda, q)| + \sum_{n=N+1}^\infty |C_n(x, \lambda, q_m) - C_n(x, \lambda, q)| \\&\leq \sum_{n=1}^N |C_n(x, \lambda, q_m) - C_n(x, \lambda, q)| + 2c\sum_{n=N+1}^\infty \frac{M^n}{n!}. \end{align*} $$ As the latter term can be made arbitrarily small by increasing $N$, it remains to show that for each $1\leq n\leq N$, the term $$ \Delta_m(x, \lambda) := C_n(x, \lambda, q_m) - C_n(x, \lambda, q) $$ can be made arbitarily small by increasing $m$ if necessary. Note that $$ \Delta_m(x, \lambda) = \left\langle p_{x,\lambda}, \prod_{i=1}^n \overline{q_m}(t_i) - \prod_{i=1}^n \overline{q}(t_i)\right\rangle_{L^2_\mathbb{C}([0,1]^n)}, $$ where $p_{x,\lambda} = c_\lambda(t_1)\prod_{i=1}^n s_\lambda (t_{i+1} - t_i)\mathbf{1}_{\{ 0\leq t_1\leq\dots\leq t_{n+1}=x \}}$.
By continuity the point $$ (x_m, \lambda_m) := \argmax_{(x,\lambda\in \overline{A}} |\Delta_m(x,\lambda)| $$ exists, and it therefore suffices to show that $|\Delta_m(x_m, \lambda_m)|\to 0$ as $m\to\infty$. So suppose otherwise for the sake of contradiction, replace $(x_m, \lambda_m)$ with a convergent subsequence so that $$(x_m, \lambda_m)\to (x_*, \lambda_*),$$ but for each $m$ we have $$|\Delta_m(x_m, \lambda_m)|\geq \delta > 0.$$
Now $p_{x_m, \lambda_m}\to p_{x_*, \lambda_*}$ by the bounded convergence theorem. Furthermore, by testing against monomials $t_1^{k_1}\dots t_n^{k_n}$, which are dense in $L^2_\mathbb{C}([0,1]^n)$, we see that $\prod_{i=1}^n q_m(t_i)\rightharpoonup \prod_{i=1}^n q(t_i)$. Therefore, we have $$ \Delta_m(x, \lambda) = \left\langle p_{x,\lambda}, \prod_{i=1}^n \overline{q_m}(t_i) - \prod_{i=1}^n \overline{q}(t_i)\right\rangle_{L^2_\mathbb{C}([0,1]^n)}\to 0, $$ which runs on contradiction with our choices of $x_m$ and $\lambda_m$. $\square$
This second-to-last part concerns the gradients and derivatives of our $y_j$. When differentiating $y_j$ with respect to $q\in L^2$, we are working with the Fréchet derivative, which yields a bounded linear map $d_qy_j:L^2_\mathbb{C}\to \mathbb{C}$. Since $L^2_\mathbb{C}$ is a Hilbert space, the Riesz representation theorem gives us a unique function, denoted $\frac{\partial y_j}{\partial q}$ in $L^2$, so that $$ d_qy_j(v) = \left\langle v, \overline{\frac{\partial y_j}{\partial q}}\right\rangle_{L^2_\mathbb{C}} \text{ for all $v\in L^2_\mathbb{C}$}. $$ The function $\frac{\partial y_j}{\partial q}$ is called the gradient of $y_j$ with respect to $q$, and it plays the analogous role of the usual gradient that appears in multivariable calculus.
To at least have an idea of how this gradient is supposed to look like, we will first proceed formally. Differentiating both sides of $$ -y_j^{\prime\prime} + q(x)y_j = \lambda y_j $$ with repsect to $q$ in the direction $v\in L^2_\mathbb{C}$, we obtain $$ -d_q(y_j^{\prime\prime})(v) + q(x)d_qy_j(v) = \lambda d_qy_j(v) - vy_j. $$ If we allow ourselves to swap the order of derivatives, we obtain $$ -(d_qy_j(v))^{\prime\prime} + q(x)d_qy_j(v) = \lambda d_qy_j(v) - vy_j, $$ and Theorem 2 lets us solve for $d_qy_j(v)$ thus: $$ d_q y_j(v) = \int_0^x y_j(t)\left[y_1(t)y_2(x) - y_1(x)y_2(t)\right]v(t)\,dt. $$ Therefore we have $$ \frac{\partial y_j}{\partial q(t)} = y_j(t)\left[y_1(t)y_2(x) - y_1(x)y_2(t)\right]\mathbf{1}_{[0,x]}(t). $$ The key obstacle in making this argument rigorous is the fact that we cannot always swap the order of derivatives. Nevertheless, one can bypass this with a density argument.
Theorem 7.
- One has $$ \frac{\partial y_j}{\partial q(t)} = y_j(t)\left[y_1(t)y_2(x) - y_1(x)y_2(t)\right]\mathbf{1}_{[0,x]}(t), $$ and $$ \frac{\partial y_j^\prime}{\partial q(t)} = y_j(t)\left[y_1(t)y_2^\prime(x) - y_1^\prime(x)y_2(t)\right]\mathbf{1}_{[0,x]}(t). $$
- One has $$ \frac{\partial y_j}{\partial \lambda} = -\int_0^1 \frac{\partial y_j}{\partial q(t)}\,dt, \qquad \frac{\partial y_j^\prime}{\partial \lambda} = -\int_0^1 \frac{\partial y_j^\prime}{\partial q(t)}\,dt. $$
Proof.
- If $q$ were continuous, then looking at the ODE we may conclude that $y_j$ is $C^2$, and so Clairaut’s theorem allows us to swap the order of derivatives. In other words, the result $$ d_q y_j(v) = \int_0^x y_j(t)\left[y_1(t)y_2(x) - y_1(x)y_2(t)\right]v(t)\,dt $$ holds for all continuous $q$. Since $C^0$ is dense in $L^2$, this result extends to all $q\in L^2$, and the conclusion follows immediately. Furthermore, if $q$ were continuous we can differentiate this integral and get $$ d_q y_j^\prime(v) = \int_0^x y_j(t)\left[y_1(t)y_2^\prime(x) - y_1^\prime(x)y_2(t)\right]v(t)\,dt, $$ from which the desired expression for $\frac{\partial y_j^\prime}{\partial q(t)}$ follows. $\square$
- Note that $y_j(x, \lambda + \varepsilon, q) = y_j(x, \lambda, q-\varepsilon)$. Therefore, we have $$ \begin{align*} \frac{\partial}{\partial \lambda}y_j(x, \lambda, q) &= \lim_{\varepsilon\to 0} \frac{y_j(x, \lambda + \varepsilon, q) - y_j(x, \lambda, q)}{\varepsilon} \\&= -\lim_{\varepsilon\to 0} \frac{y_j(x, \lambda, q-\varepsilon) - y_j(x, \lambda, q)}{-\varepsilon} \\&= -d_q y_j(1) \\&= -\int_0^1 \frac{\partial y_j}{\partial q(t)}\,dt. \end{align*} $$ The case of $y_j^\prime$ is analogously dealt with. $\square$
It is remarkable that the derivative of $y_j$ is described in terms of products of the $y_j$. This last theorem collects two useful facts to deal with products.
Theorem 8.
- For each $(\lambda, q)\in \mathbb{C}\times L^2_\mathbb{C}$ the functions $y_1^2$, $y_1y_2$ and $y_2^2$ are linearly independent over $[0,1]$.
- Let $q\in C^1_\mathbb{C}([0,1])$ and $L=q(x)\frac{d}{dx} + \frac{d}{dx}q(x) - \frac{1}{2}\left(\frac{d}{dx}\right)^3$. If $f$ and $g$ satisfy the ODE $$ -y^{\prime\prime} + q(x)y = \lambda y $$ for the same $\lambda$, then $L(fg) = 2\lambda\frac{d}{dx}(fg)$.
Proof.
- Using the ODE and the Wronskian identity gives us $$ \begin{align*} &\begin{vmatrix}y_1^2 & y_1y_2 & y_2^2 \\ (y_1^2)^{\prime} & (y_1y_2)^{\prime} & (y_2^2)^{\prime} \\ (y_1^2)^{\prime\prime} & (y_1y_2)^{\prime\prime} & (y_2^2)^{\prime\prime}\end{vmatrix} \\&=\begin{vmatrix}y_1^2 & y_1y_2 & y_2^2 \\ (y_1^2)^{\prime} & (y_1y_2)^{\prime} & (y_2^2)^{\prime} \\ (y_1^2)^{\prime\prime}-2(q-\lambda)y_1^2 & (y_1y_2)^{\prime\prime}-2(q-\lambda)y_1y_2 & (y_2^2)^{\prime\prime}-2(q-\lambda)y_2^2\end{vmatrix} \\&=\begin{vmatrix}y_1^2 & y_1y_2 & y_2^2 \\ (y_1^2)^{\prime} & (y_1y_2)^{\prime} & (y_2^2)^{\prime} \\ 2(y_1^\prime)^2 & 2y_1^\prime y_2^\prime & (y_2^\prime)^2 \end{vmatrix} \\&= 2(y_1y_2^\prime - y_1’y_2)^3 \\&= 2. \qquad\square \end{align*} $$
- One proceeds by brute-force expansion. That’s it, unfortunately. $\square$
With these two parts done, we are finally out of “prerequisite hell”! We assure the reader that the next part will be a lot more interesting than this.
Somehow the blog displays all the posts out of order. I will see if I can find a way to fix this.