\documentclass[11pt]{article} \usepackage{amssymb} \usepackage{amsfonts} \usepackage{amsmath} \usepackage{enumerate} \usepackage{bbold} \usepackage{bm} \parindent0in \pagestyle{myheadings} \DeclareSymbolFont{AMSb}{U}{msb}{m}{n} \DeclareMathSymbol{\N}{\mathbin}{AMSb}{"4E} \DeclareMathSymbol{\Z}{\mathbin}{AMSb}{"5A} \DeclareMathSymbol{\R}{\mathbin}{AMSb}{"52} \DeclareMathSymbol{\Q}{\mathbin}{AMSb}{"51} \DeclareMathSymbol{\I}{\mathbin}{AMSb}{"49} \DeclareMathSymbol{\C}{\mathbin}{AMSb}{"43} \newtheorem{theorem}{Theorem} \newenvironment{proof}[1][Proof]{\flushleft \textbf{#1.} \text }{ \hfill $\blacksquare$} \newtheorem{example0}{Example}[section] \newtheorem{lemma0}{Lemma}[section] \newtheorem{theorem0}{Theorem} \newenvironment{definition}[1][Definition]{\flushleft \textbf{#1 :}}{ } \newenvironment{example}{\medskip \flushleft \begin{example0}\rm}{\end{example0}} \newenvironment{lemma}{\medskip \flushleft \begin{lemma0}\rm}{\end{lemma0}} \renewenvironment{theorem}{\medskip \flushleft \begin{theorem0} }{\end{theorem0}} \newenvironment{Note}[1][NOTE]{\flushleft\textbf{#1 :} }{ } \renewenvironment{definition}[1][Definition]{\flushleft \textbf{#1 : }}{ } \renewenvironment{example}[1][Example]{\flushleft\textbf{#1}}{} \newenvironment{corollary}[1][Corollary :]{\flushleft \textbf{#1} }{} \newenvironment{mytheorem}[1][Theorem]{\flushleft \textbf{#1} %\textbf{} \hspace{0.05 in}\it }{ \rm } \usepackage{../Math598} \def\X{\mathbb{X}} \def\bmthzero{\bm{\theta_0}} \def\bmthone{\bm{\theta_1}} \def\bmthtwo{\bm{\theta_2}} \def\th{\theta} \def\bmphi{\bm{\phi}} \def\bmphizero{\bm{\phi_0}} \def\bmth{\bm{\theta}} \def\bmphi{\bm{\phi}} \def\bmx{\bm{x}} \def\bmg{\bm{g}} \def\transpose{{\textsf{T}}} \def\bmZn{\bm{Z}_n} \def\bmZ{\bm{Z}} \def\bmD{\bm{D}} \def\dddot#1{\stackrel{\ldots}{#1}} \def\lra{\longrightarrow} \def\Lra{\Longrightarrow} \def\Llra{\Longleftrightarrow} \def\CiP{\stackrel{p}{\lra}} \def\Cas{\stackrel{a.s.}{\lra}} \def\CiL{\stackrel{\mathfrak{L}}{\lra}} \def\eqdef{\stackrel{def}{=}} \def\tildethn{\skew2\tilde{\bmth}_n} \def\hatthn{\skew2\hat{\bmth}_n} \def\tildephin{\skew2\tilde{\bmphi}_n} \def\thnstar{\bm{\theta}_n^\star} \def\thtilde{\widetilde{\bmth}_n} \def\thhat{\widehat{\bmth}_n} \def\dotlike{\dot{\bm{l}}_n} \def\ddotlike{\ddot{\bm{l}}_n} \def\thhatNewt{\widehat{\bmth}^{(1)}} \def\thhatScor{\widehat{\bmth}^{\star}} \def\d{\textsf{d}} \def\action{\textsf{a}} \def\Loss{L} \begin{document} <>= library(knitr) # global chunk options opts_chunk$set(cache=TRUE, autodep=TRUE, fig.align='center') options(scipen=999) options(repos=c(CRAN="https://cloud.r-project.org/")) inline_hook <- function (x) { if (is.numeric(x)) { # ifelse does a vectorized comparison # If integer, print without decimal; otherwise print two places res <- ifelse(x == round(x), sprintf("%.3f", x), sprintf("%.3f", x) ) paste(res, collapse = ", ") } } #knit_hooks$set(inline = inline_hook) @ <>= myrnd <- function(x,d) {formatC(x, format="f", digits=d)} set.seed(10101) @ \begin{center} \textsclarge{MATH 559: Bayesian Theory and Methods} \vspace{0.1 in} \textsc{Basics of Decision Theory} \end{center} The key components of a \emph{decision problem} are as follows; \begin{itemize} \item a \emph{decision} $\d$ is to be made, and the decision is selected from some set $\mathcal{D}$ of alternatives. \begin{itemize} \item sometimes the terminology \emph{action} (denoted $\action$) is used. \end{itemize} \item a true \emph{state of nature}, $\upsilon(\theta)$, lying in set $\Upsilon$, defined by the data generating model, $F_Y(y;\theta)$. \item a \emph{loss function}, $\Loss(\d,\upsilon)$, for decision $\d$ and state $\upsilon$, which records the loss (or penalty) incurred when the true state of nature is $\upsilon$ and the decision made is $\d$. \end{itemize} Without loss of generality, the loss function is presumed to be non-negative. \bigskip \textbf{Decisions without data:} We can first consider the framework in the \textbf{deterministic} setting; ultimately, the decisions will be functions of data drawn from $F_Y$, with $\d \equiv \d(\by)$. We may also simplify to a discrete case, that is, where $\Upsilon$ is a finite set. This will allow us to give concrete examples of the key concepts. \bigskip \textbf{Example:} Suppose there are four possible states of nature, $\{\upsilon_1,\upsilon_2,\upsilon_3, \upsilon_4\}$, and three possible decisions $\{\d_1,\d_2,\d_3\}$. Note that precisely \textbf{one} of the states of nature is the true state (but we do not know which one). Suppose that $\Loss(\d_i,\upsilon_j) = l_{ij}$, for $i =1,2,3,4, \ j=1,2,3$, that is, the loss for making the decision $\d_i$ when the state of nature is $\upsilon_j$ is the numerical value $l_{ij}$. We may display this in a loss table \begin{table}[h] \centering \begin{tabular}{|c|c|c|c|} \hline & $\d_1$&$\d_2$&$\d_3$\\ \hline $\upsilon_1$ & \textbf{2} & 7 & 12 \\ $\upsilon_2$ & \textbf{3} & 4 & 20 \\ $\upsilon_3$ & \textbf{1} & 2 & 5 \\ $\upsilon_4$ & \textbf{5} & 6 & 9 \\ \hline \end{tabular} \end{table} For example, if the state of nature is $\upsilon_1$, and decision $\d_2$ is chosen, the loss incurred is $l_{12} = 7$. \medskip Our task is to make the decision that leads to the smallest loss. The smallest loss in each row of the table is indicated in bold. We note from the table that decision $\d_1$ has the smallest loss irrespective of the state of nature. Hence $\d_1$ is the optimal decision. \bigskip \textbf{Example:} Suppose there are three possible states of nature, $\{\upsilon_1,\upsilon_2,\upsilon_3\}$, and five possible decisions $\{\d_1,\d_2,\d_3,\d_4,\d_5\}$. We may display the numerical loss values in a loss table \begin{table}[h] \centering \begin{tabular}{|c|c|c|c|c|c|} \hline & $\d_1$&$\d_2$&$\d_3$&$\d_4$&$\d_5$\\ \hline $\upsilon_1$ & 2 & 4 & \textbf{1} & 2 & 5 \\ $\upsilon_2$ & 3 & 3 & \textbf{2} & 5 & 9 \\ $\upsilon_3$ & \textbf{1} & 5 & 4 & 2 & 4 \\ \hline \end{tabular} \end{table} From the table, we can read off the optimal decisions: \begin{itemize} \item if the true state is $\upsilon_1$, the optimal decision is $\d_3$; a loss of 1 would be incurred. \item if the true state is $\upsilon_2$, the optimal decision is $\d_3$; a loss of 2 would be incurred. \item if the true state is $\upsilon_3$, the optimal decision is $\d_1$; a loss of 1 would be incurred. \end{itemize} There is not one single optimal decision that outperforms the others for all states of nature; in general $\Loss(\d, \upsilon)$ will not have a unique minimizing $\d$ across all $\upsilon$. \medskip In this case, we would need to establish criteria via which the single optimal decision should be made. \pagebreak \textbf{Example:} Suppose there are two possible states of nature, $\{\upsilon_1,\upsilon_2\}$, and four possible decisions $\{\d_1,\d_2,\d_3,\d_4\}$. Suppose that the loss table takes the form \begin{table}[h] \centering \begin{tabular}{|c|c|c|c|c|} \hline & $\d_1$&$\d_2$&$\d_3$&$\d_4$\\ \hline $\upsilon_1$ & 5 & \textbf{2} & 9 & 12 \\ $\upsilon_2$ & \textbf{4} & 8 & 5 & 7 \\ \hline \end{tabular} \end{table} \begin{itemize} \item if the true state is $\upsilon_1$, the optimal decision is $\d_2$; a loss of 2 would be incurred. \item if the true state is $\upsilon_2$, the optimal decision is $\d_1$; a loss of 4 would be incurred. \end{itemize} We can depict this loss table in a plot using the columns as coordinates in $(\upsilon_1,\upsilon_2)$ space. <>= u1.loss<-c(5,2,9,12); u2.loss<-c(4,8,5,7) par(mar=c(5,4,1,2)) plot(u1.loss,u2.loss,pch=19,xlim=range(0,12),ylim=range(0,12), xlab=expression(paste('Loss under ',upsilon[1])), ylab=expression(paste('Loss under ',upsilon[2]))) text(u1.loss,u2.loss-0.75, c(expression(d[1]),expression(d[2]),expression(d[3]),expression(d[4]))) @ One strategy for choosing the decision would be to choose the one leads to the smallest possible loss, which would be $\d_2$, incurred when $\upsilon_1$ occurs. However, this would lead to potential difficulties, as the loss for this decision if $\upsilon_2$ occurs is 8, which is large. \medskip This plot does clarify, however, that certain decisions should not be considered; $\d_3$ and $\d_4$ are inferior decisions irrespective of whether $\upsilon_1$ or $\upsilon_2$ is the true state. This encourages us to consider the `lower left' boundary of the convex hull of the points. <>= par(mar=c(5,4,1,2)) plot(u1.loss,u2.loss,pch=19,xlim=range(0,12),ylim=range(0,12), xlab=expression(paste('Loss under ',upsilon[1])), ylab=expression(paste('Loss under ',upsilon[2]))) text(u1.loss,u2.loss-0.75, c(expression(d[1]),expression(d[2]),expression(d[3]),expression(d[4]))) polygon(u1.loss[c(1,2,4,3)],u2.loss[c(1,2,4,3)],col='cyan') @ \textbf{Admissibility:} A decision rule $\d$ is \textit{inadmissible} if there exists another rule $\d^\ast$ such that \[ R_n(\d^\ast,\upsilon) \leq R_n(\d,\upsilon) \quad \text{for all } \upsilon \in \Upsilon \] and \[ R_n(\d^\ast,\upsilon) < R_n(\d,\upsilon) \quad \text{for at least one } \upsilon \in \Upsilon \] that is, the decision rule $\d$ yields frequentist risk that is \textbf{at least as large} as the frequentist risk for $\d^\ast$, for all possible states of nature, and is \textbf{greater} than the frequentist risk for $\d^\ast$ for at least one value of $\upsilon$. If $\d$ is not inadmissible, then it is \textit{admissible}. \bigskip \textbf{Randomized decisions:} Consider the line segment connecting the $\d_1$ and $\d_2$ points, that is, the line constructed as \[ p \d_1 + (1-p) \d_2 \qquad 0 \leq p \leq 1 \] We can consider making the following \emph{randomized} decision: for some $p$, we pick decision $\d_1$ with probability $p$, and decision $\d_2$ with probability $(1-p)$. In this case, our expected loss will be \begin{align*} \upsilon_1 \text{ true} & : \ p \Loss(\d_1,\upsilon_1) + (1-p) \Loss(\d_2,\upsilon_1) = 5 p+ 2(1-p) = 3p+2\\[6pt] \upsilon_2 \text{ true} & : \ p \Loss(\d_1,\upsilon_2) + (1-p) \Loss(\d_2,\upsilon_2) = 4 p+ 8(1-p) = -4p + 8 \end{align*} These two losses are equal if $p=6/7$, when the loss is $32/7$. The randomized decision approach provides an alternative strategy to choosing a decision; we can generalize this to consider an arbitrary convex combination of all four decisions, and consider three probabilities $(p_1,p_2,p_3)$ whose sum is not greater than 1, with $p_4 = (1-p_1 - p_2 - p_3)$, then with probability $p_j$ choose decision $\d_j$. The randomized decision then falls within the \emph{convex hull} of the four points. Plotting the location of the $p=6/7$ randomized decision and the convex hull, we obtain the following plot: <>= par(mar=c(4,4,1,2)) plot(u1.loss,u2.loss,pch=19,xlim=range(0,12),ylim=range(0,12), xlab=expression(paste('Loss under ',upsilon[1])), ylab=expression(paste('Loss under ',upsilon[2]))) text(u1.loss,u2.loss-0.75, c(expression(d[1]),expression(d[2]),expression(d[3]),expression(d[4]))) slope<-(u2.loss[2]-u2.loss[1])/(u1.loss[2]-u1.loss[1]) intercept<-u2.loss[2]-slope*u1.loss[2] abline(intercept,slope,lty=3) polygon(u1.loss[c(1,2,4,3)],u2.loss[c(1,2,4,3)],col='cyan') legend(0,12,c('p=6/7'),col='red',pch=15) points((6/7)*u1.loss[1]+(1/7)*u1.loss[2], (6/7)*u2.loss[1]+(1/7)*u2.loss[2],pch=15,col='red') @ \pagebreak \textbf{Bayes decisions:} Bayes decision rules posit a prior distribution over the space of states of nature, and computes the expected frequentist risk over that distribution. That is, in the discrete case, the Bayes risk associated with decision $\d$ is \[ R(\d) = \E_{\pi_0}[\Loss(\d,\upsilon)] = \sum_{\upsilon \in \Upsilon} \Loss(\d,\upsilon) \pi_0(\upsilon) \] and we can the compare the decisions directly in terms of this expected loss. In the continuous case we have \[ R(\d) = \E_{\pi_0}[\Loss(\d,\upsilon)] = \int \Loss(\d,\upsilon) \pi_0(\upsilon) \ d \upsilon. \] To make it explicit that this expected risk is prior-dependent, we may write $R^{\pi_0}(\d)$. \bigskip \textbf{Example:} For the loss table above \begin{table}[h] \centering \begin{tabular}{|c|c|c|c|c|} \hline & $\d_1$&$\d_2$&$\d_3$&$\d_4$\\ \hline $\upsilon_1$ & 5 & \textbf{2} & 9 & 12 \\ $\upsilon_2$ & \textbf{4} & 8 & 5 & 7 \\ \hline \end{tabular} \end{table} we can report the Bayes risk for each decision according to a prior that places $(\pi_{01},1-\pi_{01})$ on the two states of nature. If the probabilities are $(0.2,0.8)$, we have the result \begin{table}[h] \centering \begin{tabular}{|c|c|c|c|c|c|} \hline & $\pi_0$ & $\d_1$&$\d_2$&$\d_3$&$\d_4$\\ \hline $\upsilon_1$ & 0.2 & 5 & \textbf{2} & 9 & 12 \\ $\upsilon_2$ & 0.8 & \textbf{4} & 8 & 5 & 7 \\ \hline Bayes risk & & \textbf{\Sexpr{myrnd(0.2*u1.loss[1]+0.8*u2.loss[1],1)}} & \Sexpr{myrnd(0.2*u1.loss[2]+0.8*u2.loss[2],1)} & \Sexpr{myrnd(0.2*u1.loss[3]+0.8*u2.loss[3],1)} & \Sexpr{myrnd(0.2*u1.loss[4]+0.8*u2.loss[4],1)} \\ \hline \end{tabular} \end{table} \noindent whereas if the probabilities are $(0.9,0.1)$, we have the result \begin{table}[h] \centering \begin{tabular}{|c|c|c|c|c|c|} \hline & $\pi_0$ & $\d_1$&$\d_2$&$\d_3$&$\d_4$\\ \hline $\upsilon_1$ & 0.9 & 5 & \textbf{2} & 9 & 12 \\ $\upsilon_2$ & 0.1 & \textbf{4} & 8 & 5 & 7 \\ \hline Bayes risk & &\Sexpr{myrnd(0.9*u1.loss[1]+0.1*u2.loss[1],1)} & \textbf{\Sexpr{myrnd(0.9*u1.loss[2]+0.1*u2.loss[2],1)}} & \Sexpr{myrnd(0.9*u1.loss[3]+0.1*u2.loss[3],1)} & \Sexpr{myrnd(0.9*u1.loss[4]+0.1*u2.loss[4],1)} \\ \hline \end{tabular} \end{table} We can plot a prior as lines on the plot above by considering points of equal risk, that is, where \[ \pi_{01} L_1 + (1-\pi_{01}) L_2 = r \] say. In the following plot, we give the lines \[ y = \frac{r}{1-\pi_{01}}-\frac{\pi_{01}}{1-\pi_{01}}x \] for different values of $r$, and $\pi_{01}=0.2$. <>= par(mar=c(4,4,1,2)) plot(u1.loss,u2.loss,pch=19,xlim=range(0,12),ylim=range(0,12), xlab=expression(paste('Loss under ',upsilon[1])), ylab=expression(paste('Loss under ',upsilon[2]))) text(u1.loss,u2.loss-0.75, c(expression(d[1]),expression(d[2]),expression(d[3]),expression(d[4]))) polygon(u1.loss[c(1,2,4,3)],u2.loss[c(1,2,4,3)],col='cyan') pi.01<-0.2 r<-1;abline(r/(1-pi.01),-pi.01/(1-pi.01),col='blue',lty=3) r<-2;abline(r/(1-pi.01),-pi.01/(1-pi.01),col='blue',lty=3) r<-3;abline(r/(1-pi.01),-pi.01/(1-pi.01),col='blue',lty=3) r<-4;abline(r/(1-pi.01),-pi.01/(1-pi.01),col='blue',lty=3) legend(0,12,c('Lines of constant risk'),lty=3,col='blue') @ It is evident that for this prior, the optimal decision will be where the blue dotted lines touch the convex hull for the first time, when the decision will be $\d_1$, and the risk will be 4.2. Furthermore, $\d_1$ will be optimal decision for \textbf{all} priors such that the slope $-\pi_{01}/(1-\pi_{01})$ is less negative than the slope of the line connecting $\d_1$ to $\d_2$; this slope is $\Sexpr{u2.loss[2]-u2.loss[1]}/(\Sexpr{u1.loss[2]-u1.loss[1]}) = \Sexpr{myrnd(slope,3)}$, so we need \[ \frac{\pi_{01}}{1-\pi_{01}} > \Sexpr{myrnd(-slope,3)} \qquad \Longrightarrow \qquad \pi_{01} > \frac{4}{7}. \] If $\pi_{01} < 4/7$, the optimal decision is $\d_2$, and the risk will be 4.9. If $\pi_{01} = 4/7$, the two decisions give equal risk and the dotted blue line coincides with the line from $\d_1$ to $\d_2$: \Sexpr{pi.01<-4/7} \begin{table}[h] \centering \begin{tabular}{|c|c|c|c|c|c|} \hline & $\pi_0$ & $\d_1$&$\d_2$&$\d_3$&$\d_4$\\ \hline $\upsilon_1$ & 4/7 & 5 & \textbf{2} & 9 & 12 \\ $\upsilon_2$ & 3/7 & \textbf{4} & 8 & 5 & 7 \\ \hline Bayes risk & & \textbf{\Sexpr{myrnd(pi.01*u1.loss[1]+(1-pi.01)*u2.loss[1],1)}} & \textbf{\Sexpr{myrnd(pi.01*u1.loss[2]+(1-pi.01)*u2.loss[2],1)}} & \Sexpr{myrnd(pi.01*u1.loss[3]+(1-pi.01)*u2.loss[3],1)} & \Sexpr{myrnd(pi.01*u1.loss[4]+(1-pi.01)*u2.loss[4],1)} \\ \hline \end{tabular} \end{table} \bigskip The \emph{least favourable prior}, $\pi_0^\dag$, is the prior for which the expected loss is as large as possible, that is, the prior for which \[ R^{\pi_0}(\d) \equiv \E_{\pi_0}[\Loss(\d,\upsilon)] \leq \E_{\pi_0^\dag}[\Loss(\d,\upsilon)] \equiv R^{\pi_0^\dag}(\d) \] Note that here the least favourable prior is decision-dependent. For the example above, the expected loss for decision $\d_1$ is \[ \pi_{01} \times \Sexpr{u1.loss[1]} + (1-\pi_{01}) \Sexpr{u2.loss[1]} = \Sexpr{u2.loss[1]} + \pi_{01} \times \Sexpr{u1.loss[1]-u2.loss[1]} \] so the least favourable prior is when $\pi_{01}$ is as \textbf{large} as possible, that is, when $\pi_{01} = 4/7$. For decision $\d_2$, the expected loss is \[ \pi_{01} \times \Sexpr{u1.loss[2]} + (1-\pi_{01}) \Sexpr{u2.loss[2]} = \Sexpr{u2.loss[2]} + \pi_{01} \times (\Sexpr{u1.loss[2]-u2.loss[2]}) \] so the least favourable prior is when $\pi_{01}$ is as \textbf{small} as possible, that is, again when $\pi_{01} = 4/7$. \pagebreak \textbf{Minimax decisions:} The \textit{minimax} decision is the one that \textbf{minimizes} the \textbf{maximum} frequentist risk after all states of nature have been considered. That is \[ \widehat \d_M = \min_{\d} \max_\upsilon R(\d,\upsilon). \] More formally, to allow for the general nature of the two optimizations, we should write \[ \widehat \d_M = \inf_{\d} \sup_\upsilon R(\d,\upsilon). \] although in most cases the first version is adequate. \bigskip \textbf{Example:} For the loss table \begin{table}[h!] \centering \begin{tabular}{|c|c|c|c|c|} \hline & $\d_1$&$\d_2$&$\d_3$&$\d_4$\\ \hline $\upsilon_1$ & 5 & 2 & 9 & 12 \\ $\upsilon_2$ & 4 & 8 & 5 & 7 \\ \hline \text{Maximum} & 5 & 8 & 9 & 12 \\ \hline \end{tabular} \end{table} the minimax decision is $\d_1$, incurring losses 5 and 4 under the two states of nature. \bigskip \textbf{Example:} Consider the modified loss table \begin{table}[h!] \centering \begin{tabular}{|c|c|c|c|c|} \hline & $\d_1$&$\d_2$&$\d_3$&$\d_4$\\ \hline $\upsilon_1$ & 3 & 8&5&20 \\ $\upsilon_2$ & 5 & 1&2&20 \\ \hline \text{Maximum} & 5 & 8 & 5 & 20 \\ \hline \end{tabular} \end{table} In this case, the (pure) minimax decisions are $\d_1$ and $\d_3$ with minimax risk 5. <>= u1.loss<-c(3,8,5,20); u2.loss<-c(5,1,2,20) par(mar=c(5,4,1,2)) plot(u1.loss,u2.loss,pch=19,xlim=range(0,20),ylim=range(0,20), xlab=expression(paste('Loss under ',upsilon[1])), ylab=expression(paste('Loss under ',upsilon[2]))) text(u1.loss,u2.loss-0.75, c(expression(d[1]),expression(d[2]),expression(d[3]),expression(d[4]))) polygon(u1.loss[c(1,3,2,4)],u2.loss[c(1,3,2,4)],col='cyan') @ Now consider the randomized decision with probability $p$: this yields loss \begin{align*} \upsilon_1 \text{ true} & : \ p \Loss(\d_1,\upsilon_1) + (1-p) \Loss(\d_2,\upsilon_1) = 3 p+ 5(1-p) = -2p+5\\[6pt] \upsilon_2 \text{ true} & : \ p \Loss(\d_1,\upsilon_2) + (1-p) \Loss(\d_2,\upsilon_2) = 5 p+ 2(1-p) = 3p+2 \end{align*} and the two losses are equal when $5p=3$, so $p=3/5$, that is, when the loss is equal to 19/5. This loss is \textbf{lower} than the non-randomized minimax risk. Furthermore, the randomized decision itself a minimax decision, as depicted in the next plot. <>= par(mar=c(4,4,2,2)) plot(0.5,0.5,xlim=range(0,1),ylim=range(0,6), type='n', xlab=expression(p),ylab='Loss') abline(5,-2); abline(2,3,lty=3); lines(c(0,0.6,1),c(5,3.8,5),col='red') legend(0,1.8,c(expression(paste('Loss under ',upsilon[1])), expression(paste('Loss under ',upsilon[2])), 'Maximum loss'),lty=c(1,2,1),col=c('black','black','red')) @ The red line indicates the maximum loss across the two states of nature as $p$ varies; the minimum loss achievable is 19/5 = 3.8, whereas the pure minimax loss is 5. \pagebreak \textbf{Statistical Decision Theory:} We now acknowledge that the decisions may be functions of observed data. The true state of nature will be some function of the data generating model parameter $\theta$, so for convenience we will write $\upsilon(\theta) \equiv \theta$, and note that the true state of nature will in general lie in some region of $\R^d$ that is not simply a finite set. Our task is to identify the optimal decision function, $\d(\cdot)$. \begin{itemize} \item \textbf{Frequentist risk:} The \emph{frequentist risk} or \emph{loss} associated with decision denoted $\d \equiv \d(\bY)$ is the expected loss associated with $\d (\bY)$, with the expectation taken over the distribution of $\bY$ given $\theta$: \begin{equation*} R_{n} ({\d,\theta}) = \E_{F_{\bY}} [ \Loss(\d(\bY),\theta ) ; \theta] = \int_\mathcal{Y} \Loss(\d(\by),\theta )f_{\bY}(\by;\theta ) \: d\by \end{equation*} \item \textbf{Bayes risk:} The (\emph{average}) \emph{Bayes risk} for $\d(\bY)$ is the expected frequentist risk over \emph{prior} distribution, $\pi_0(\theta)$, of $\theta$: in the continuous case \begin{align*} R_n(\d ) = \E_{\pi_0 }[ R_{n} ({\d,\theta}) ] & = \E_{\pi_0 } \left[ \E_{F_\bY}\left[ \Loss(\d(\bY),\theta ); \theta \right] \right] & \textrm{iterated expectation}\\[6pt] & = \int_\Theta \left\{ \int_\mathcal{Y} \Loss(\d(\by),\theta )f_{\bY}(\by;\theta ) \: d \by \right\} \pi_{0}(\theta)\: d\theta \\[6pt] & = \int_\mathcal{Y} \left\{ \int_\Theta \Loss(\d(\by),\theta )\pi_{n}(\theta)\: d\theta \right\} f_{\bY}(\by) \: d \by & \textrm{Bayes theorem} \end{align*} where the Bayes theorem identity is $\pi_n(\theta) = f_{\bY}(\by;\theta ) \pi_0(\theta)/f_{\bY}(\by)$. We can see that if we choose the decision mechanism $\d (\cdot)$ such that the inner integral \[ \int_\Theta \Loss(\d(\by),\theta )\pi_{n}(\theta)\: d\theta \] is minimized for each $\by$, then this decision will \textbf{necessarily} minimize $R_n(\d )$. That is, we seek to find the decision function to minimize the posterior expected loss; note that this decision is optimal with respect to the chosen prior $\pi_0(\theta)$. \end{itemize} \medskip \textbf{Note:} In the computation of the Bayes risk, it is assumed that $\pi_0$ is \textit{proper} (that is, it integrates to one), however the calculation may be performed with an \textit{improper} prior (that is, a prior that is not integrable) also, in which case the terminology \textit{generalized Bayes risk} is used. It is sometimes useful to consider a sequence of proper priors whose limiting form is improper: for example if $\pi_0(\theta) \equiv Normal(0,\lambda_0^{-1})$, for $\lambda_0 > 0$, this prior is proper, but it becomes improper in the limit as $\lambda_0 \longrightarrow 0$. \bigskip \textbf{Example:} If $f_Y(y;\theta) \equiv Normal(\mu,\sigma^2)$, then the sample mean $\overline{Y}$ is the limit of a sequence of Bayes estimators under quadratic loss if the sequence of priors is $Normal(0,\lambda_0^{-1})$ with $\lambda_0 \longrightarrow 0$ \bigskip It can be shown that \textit{Bayes rules} (the decisions that minimize the Bayes risk) have desirable properties: \begin{mytheorem} Every Bayes decision made with respect to prior $\pi_0$ such that $\pi_0(\theta) > 0$ for all $ \theta \in \Theta$, and such that $R_n(d) < \infty$, is admissible. \end{mytheorem} \begin{proof} By contradiction: suppose $\d \equiv \d(\by)$ is the Bayes decision, but is inadmissible. Then there exists another decision $\d^\ast$ such that we have \[ \Loss(\d^\ast,\theta) \leq \Loss(\d,\theta) \quad \forall \theta. \] Multiplying both sides by $\pi_0(\theta)$ and integrating both sides, we have that \[ R_n(\d^\ast) \leq R_n(\d) < \infty \] by assumption. By assumption of inadmissibility, there must exist a $\theta$ such that $\Loss(\d^\ast,\theta) < \Loss(\d,\theta)$ for that $\theta$. But by assumption this value receives positive prior density, so in fact we must have that $R_n(\d^\ast) < R_n(\d)$. This contradicts the assumption that $\d$ is the Bayes decision. \end{proof} \pagebreak This result can be augmented to confirm the desirable features of Bayes decision rules: for example, we can show that \begin{itemize} \item Every admissible decision rule is a (generalized) Bayes decision rule for some prior $\pi_0$ (which may be improper). \item If every risk function is continuous, then every Bayes procedure with finite Bayes risk (for prior with positive density for all $\R$) is admissible. \end{itemize} \textbf{Connection between Bayes and minimax:} \begin{mytheorem} Suppose that prior $\pi_0$ is proper, and denote by $\d^{\pi_0}$ the Bayes decision rule computed with respect to $\pi_0$. Then the Bayes risk for $\d^{\pi_0}$ is no greater than the minimax risk. \end{mytheorem} \begin{proof} We have for the Bayes risk for $\d^{\pi_0}$ \begin{align*} R_n(\d^{\pi_0}) = \inf_\d R_n(d) & = \inf_\d \int R_n(\d,\theta) \pi_0(\theta) \ d \theta \\ & \leq \inf_\d \int \sup_t R_n(\d,t) \pi_0(\theta) \ d \theta \\ & = \inf_\d \sup_\theta R_n(\d,\theta) \end{align*} \end{proof} \begin{mytheorem} Suppose that prior $\pi_0$ is proper, and denote by $\d^{\pi_0}$ the Bayes decision rule computed with respect to $\pi_0$. Suppose that $\d^{\pi_0}$ has constant risk on $\Theta$. Then $\d^{\pi_0}$ is minimax, and if $\d^{\pi_0}$ is the unique Bayes decision, it is the unique minimax decision. \end{mytheorem} \begin{proof} Suppose that $\d$ is any other decision. Then \begin{align*} \sup_\theta \R_n(\d,\theta) & = \int \sup_t \R_n(\d,t) \pi_0(\theta) \ d \theta \\ & \geq \int \R_n(\d,\theta) \pi_0(\theta) \ d \theta \\ & \geq \int \R_n(\d^{\pi_0},\theta) \pi_0(\theta) \ d \theta & \text{as $\d^{\pi_0}$ is Bayes}\\ & = \sup_\theta \R_n(\d^{\pi_0},\theta) & \text{as $\d^{\pi_0}$ has constant risk} \end{align*} If $\d^{\pi_0}$ is the unique Bayes decision, then the second inequality is strict, and the minimax decision is unique. \end{proof} \begin{itemize} \item A unique minimax procedure is admissible. \item The minimax procedure has constant risk across $\Theta$. \item An admissible minimax procedure is Bayes for some prior; its risk is constant on the set of $\R$ for which the prior density is positive. \item The prior corresponding to an admissible minimax procedure is the least favourable prior. \end{itemize} \medskip \textbf{References:} These texts give much more detail of statistical decision theory: \begin{itemize} \item Robert, CP (2004), \textit{The Bayesian Choice}, Chapter 2. \item Lehmann, EL and Casella, G, (1998), \textit{Theory of Point Estimation}, Chapters 4-6. \item Shao, J (2003), \textit{Mathematical Statistics}, Chapters 2-4. \end{itemize} \end{document}