1.3: Differentiation on the computer
- Page ID
- 218598
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\( \newcommand{\dsum}{\displaystyle\sum\limits} \)
\( \newcommand{\dint}{\displaystyle\int\limits} \)
\( \newcommand{\dlim}{\displaystyle\lim\limits} \)
\( \newcommand{\id}{\mathrm{id}}\) \( \newcommand{\Span}{\mathrm{span}}\)
( \newcommand{\kernel}{\mathrm{null}\,}\) \( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\) \( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\) \( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\id}{\mathrm{id}}\)
\( \newcommand{\Span}{\mathrm{span}}\)
\( \newcommand{\kernel}{\mathrm{null}\,}\)
\( \newcommand{\range}{\mathrm{range}\,}\)
\( \newcommand{\RealPart}{\mathrm{Re}}\)
\( \newcommand{\ImaginaryPart}{\mathrm{Im}}\)
\( \newcommand{\Argument}{\mathrm{Arg}}\)
\( \newcommand{\norm}[1]{\| #1 \|}\)
\( \newcommand{\inner}[2]{\langle #1, #2 \rangle}\)
\( \newcommand{\Span}{\mathrm{span}}\) \( \newcommand{\AA}{\unicode[.8,0]{x212B}}\)
\( \newcommand{\vectorA}[1]{\vec{#1}} % arrow\)
\( \newcommand{\vectorAt}[1]{\vec{\text{#1}}} % arrow\)
\( \newcommand{\vectorB}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\( \newcommand{\vectorC}[1]{\textbf{#1}} \)
\( \newcommand{\vectorD}[1]{\overrightarrow{#1}} \)
\( \newcommand{\vectorDt}[1]{\overrightarrow{\text{#1}}} \)
\( \newcommand{\vectE}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash{\mathbf {#1}}}} \)
\( \newcommand{\vecs}[1]{\overset { \scriptstyle \rightharpoonup} {\mathbf{#1}} } \)
\(\newcommand{\longvect}{\overrightarrow}\)
\( \newcommand{\vecd}[1]{\overset{-\!-\!\rightharpoonup}{\vphantom{a}\smash {#1}}} \)
\(\newcommand{\avec}{\mathbf a}\) \(\newcommand{\bvec}{\mathbf b}\) \(\newcommand{\cvec}{\mathbf c}\) \(\newcommand{\dvec}{\mathbf d}\) \(\newcommand{\dtil}{\widetilde{\mathbf d}}\) \(\newcommand{\evec}{\mathbf e}\) \(\newcommand{\fvec}{\mathbf f}\) \(\newcommand{\nvec}{\mathbf n}\) \(\newcommand{\pvec}{\mathbf p}\) \(\newcommand{\qvec}{\mathbf q}\) \(\newcommand{\svec}{\mathbf s}\) \(\newcommand{\tvec}{\mathbf t}\) \(\newcommand{\uvec}{\mathbf u}\) \(\newcommand{\vvec}{\mathbf v}\) \(\newcommand{\wvec}{\mathbf w}\) \(\newcommand{\xvec}{\mathbf x}\) \(\newcommand{\yvec}{\mathbf y}\) \(\newcommand{\zvec}{\mathbf z}\) \(\newcommand{\rvec}{\mathbf r}\) \(\newcommand{\mvec}{\mathbf m}\) \(\newcommand{\zerovec}{\mathbf 0}\) \(\newcommand{\onevec}{\mathbf 1}\) \(\newcommand{\real}{\mathbb R}\) \(\newcommand{\twovec}[2]{\left[\begin{array}{r}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\ctwovec}[2]{\left[\begin{array}{c}#1 \\ #2 \end{array}\right]}\) \(\newcommand{\threevec}[3]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\cthreevec}[3]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \end{array}\right]}\) \(\newcommand{\fourvec}[4]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\cfourvec}[4]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \end{array}\right]}\) \(\newcommand{\fivevec}[5]{\left[\begin{array}{r}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\cfivevec}[5]{\left[\begin{array}{c}#1 \\ #2 \\ #3 \\ #4 \\ #5 \\ \end{array}\right]}\) \(\newcommand{\mattwo}[4]{\left[\begin{array}{rr}#1 \amp #2 \\ #3 \amp #4 \\ \end{array}\right]}\) \(\newcommand{\laspan}[1]{\text{Span}\{#1\}}\) \(\newcommand{\bcal}{\cal B}\) \(\newcommand{\ccal}{\cal C}\) \(\newcommand{\scal}{\cal S}\) \(\newcommand{\wcal}{\cal W}\) \(\newcommand{\ecal}{\cal E}\) \(\newcommand{\coords}[2]{\left\{#1\right\}_{#2}}\) \(\newcommand{\gray}[1]{\color{gray}{#1}}\) \(\newcommand{\lgray}[1]{\color{lightgray}{#1}}\) \(\newcommand{\rank}{\operatorname{rank}}\) \(\newcommand{\row}{\text{Row}}\) \(\newcommand{\col}{\text{Col}}\) \(\renewcommand{\row}{\text{Row}}\) \(\newcommand{\nul}{\text{Nul}}\) \(\newcommand{\var}{\text{Var}}\) \(\newcommand{\corr}{\text{corr}}\) \(\newcommand{\len}[1]{\left|#1\right|}\) \(\newcommand{\bbar}{\overline{\bvec}}\) \(\newcommand{\bhat}{\widehat{\bvec}}\) \(\newcommand{\bperp}{\bvec^\perp}\) \(\newcommand{\xhat}{\widehat{\xvec}}\) \(\newcommand{\vhat}{\widehat{\vvec}}\) \(\newcommand{\uhat}{\widehat{\uvec}}\) \(\newcommand{\what}{\widehat{\wvec}}\) \(\newcommand{\Sighat}{\widehat{\Sigma}}\) \(\newcommand{\lt}{<}\) \(\newcommand{\gt}{>}\) \(\newcommand{\amp}{&}\) \(\definecolor{fillinmathshade}{gray}{0.9}\)Next before diving into solving (eq:1.1) we need to review how differentiation is done on the computer. To start, recall the definition of the derivative of \(y(t)\) which you probably learned in high school:
\[ \frac{d y}{d t} = \lim_{{\Delta t}\to 0} \frac{y(t + \Delta t) - y(t)}{\Delta t} \label{eq:1.5}\]
In doing numerical computation, we can’t represent the limiting transition \(\lim_{{\Delta t}\to 0}\) – our data is usually discrete (e.g. a sampled function). More importantly, floating point numbers don’t smoothly transition to zero – there is no "infinitesimally small but nonzero" number representable on a computer. Instead, we use the definition \ref{eq:1.5} to form an approximation to the derivative as
\[ \frac{d y}{d t} \approx \frac{y(t + h) - y(t)}{h} \label{eq:1.6}\]
where \(h\) is a small number. This is called the "forward difference" approximation of the derivative.
First derivatives from Taylor’s series
It’s easy to write down the approximation \ref{eq:1.6} by remembering the high school definition of a derivative. However, we gain some appreciation for the accuracy of approximating derivatives by deriving them using Taylor series expansions. Suppose we know the value of a function \(y\) at the time \(t\) and want to know its value at a later time \(t+h\). How to find \(y(t+h)\)? The Taylor’s series expansion says we can find the value at the later time by
\[ y(t+h) = y(t) + h \frac{d y}{d t} \biggr\rvert_{t} + \frac{h^2}{2} \frac{d^2 y}{d t^2} \biggr\rvert_{t} + \frac{h^3}{6} \frac{d^3 y}{d t^3} \biggr\rvert_{t} + O(h^4) \label{eq:1.7}\]
The notation \(O(h^4)\) means that there are terms of power \(h^4\) and above in the sum, but we will consider them details which we can ignore.
We can rearrange \ref{eq:1.7} to put the first derivative on the LHS, yielding
\[\nonumber \frac{d y}{d t} \biggr\rvert_{t} = \frac{y(t+h) - y(t)}{h} - \frac{h}{2} \frac{d^2 y}{d t^2} \biggr\rvert_{t} - \frac{h^2}{6} \frac{d^3 y}{d t^3} \biggr\rvert_{t} - O(h^4)\]
Now we make the assumption that \(h\) is very small, so we can throw away terms of order \(h\) and above from this expression. With that assumption we get
\[\nonumber y(t+h) \approx y(t) + h \frac{d y}{d t} \biggr\rvert_{t}\]
which can be rearranged to give
\[ \frac{d y}{d t} \approx \frac{y(t + h) - y(t)}{h} \label{eq:1.8}\]
i.e. the same forward difference expression as found in \ref{eq:1.6} earlier. However, note that to get this approximation we threw away all terms of order \(h\) and above. This implies that when we use this approximation we incur an error! That error will be on the order of \(h\). In the limit \(h \rightarrow 0\) the error disappears but since computers can only use finite \(h\), there will always be an error in our computations, and the error is proportional to \(h\). We express this concept by saying the error scales as \(O(h)\).
Using Taylor’s series expansions we can discover other results. Consider sitting at the time point \(t\) and asking for the value of \(y\) at time \(t - h\) in the past. In this case, we can write an expression for the past value of \(y\) as
\[ y(t-h) = y(t) - h \frac{d y}{d t} \biggr\rvert_{t} + \frac{h^2}{2} \frac{d^2 y}{d t^2} \biggr\rvert_{t} - \frac{h^3}{6} \frac{d^3 y}{d t^3} \biggr\rvert_{t} + O(h^4) \label{eq:1.9}\]
Note the negative signs on the odd-order terms. Similar to the forward difference expression above, it’s easy to rearrange this expression and then throw away terms of \(O(h)\) and above to get a different approximation for the derivative,
\[ \frac{d y}{d t} \approx \frac{y(t) - y(t-h)}{h} \label{eq:1.10}\]
This is called the "backward difference" approximation to the derivative. Since we threw away terms of \(O(h)\) and above, this expression for the derivative also carries an error penalty of \(O(h)\).
Now consider forming difference between Taylor series \ref{eq:1.7} and \ref{eq:1.9}. When we subtract one from the other we get
\[ \nonumber y(t+h) - y(t-h) = 2 h \frac{d y}{d t} \biggr\rvert_{t} + 2 \frac{h^3}{6} \frac{d^3 y}{d t^3} \biggr\rvert_{t} + O(h^5) \label{eq:TaylorsExpansionDifference}\]
This expression may be rearranged to become
\[\nonumber \frac{d y}{d t} \biggr\rvert_{t} = \frac{y(t+h) - y(t-h)}{2 h} - \frac{h^2}{6} \frac{d^3 y}{d t^3} \biggr\rvert_{t} + O(h^4)\]
Now if we throw away terms of \(O(h^2)\) and above, we get the so-called "central difference" approximation for the derivative,
\[ \frac{d y}{d t} \biggr\rvert_{t} \approx \frac{y(t+h) - y(t-h)}{2 h} \label{eq:1.11}\]
Note that to get this expression we discarded \(O(h^2)\) terms. This implies that the error incurred upon using this expression for a derivative is \(O(h^2)\). Since we usually have \(h \ll 1\), the central difference error \(O(h^2)\) is smaller than the \(O(h)\) incurred when using either the forward or backward difference formulas. That is, the central difference expression \ref{eq:1.11} provides a more accurate approximation to the derivative than either the forward or the backward difference formulas, and should be your preferred derivative approximation when you can use it.
In each of these derivations we threw away terms of some order \(p\) (and above). Discarding terms of order \(p\) and above means we no longer have a complete Taylor series. This operation is called "truncation", meaning we have truncated (chopped off) the remaining terms of the full Taylor series. The error incurred by this operation is therefore called "truncation error".
Second derivatives from Taylor’s series
We can also get an approximation for the second derivative from expressions \ref{eq:1.7} and \ref{eq:1.9}. This time, form the sum of the two expressions to get
\[\nonumber \label{eq:TaylorsExpansionSum} y(t+h) + y(t-h) = 2 y(t) + 2 \frac{h^2}{2} \frac{d^2 y}{d t^2} \biggr\rvert_{t} +O(h^4)\]
This may be rearranged to get an expression for the second derivative,
\[\nonumber \frac{d^2 y}{d t^2} \biggr\rvert_{t} = \frac{y(t+h) -2 y(t) + y(t-h)} {h^2} -O(h^2)\]
If we drop the \(O(h^2)\) term, we get the approximation to the second derivative,
\[ \frac{d^2 y}{d t^2} \biggr\rvert_{t} \approx \frac{y(t+h) -2 y(t) + y(t-h)} {h^2} \label{eq:1.12}\]
And since we dropped the \(O(h^2)\) term to make this approximation, it means that use of \ref{eq:1.12} will incur an error of \(O(h^2)\) when using this approximation.

