# Prediction or Criticism of the model mean()

**URL:** <https://discourse.edwardlib.org/t/prediction-or-criticism-of-the-model-mean/314>\
**Category:** General\
**Created:** [August 24, 2017, 2:06am UTC](https://discourse.edwardlib.org/t/prediction-or-criticism-of-the-model-mean/314 "2017-08-24T02:06:42Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![MushroomHunting](https://avatars.discourse-cdn.com/v4/letter/m/90ced4/32.png) [@MushroomHunting](https://discourse.edwardlib.org/u/MushroomHunting)\
**Post date:** [August 24, 2017, 2:06am UTC](https://discourse.edwardlib.org/t/prediction-or-criticism-of-the-model-mean/314/1 "2017-08-24T02:06:43Z")

</div>

Hi all

I’m trying to run criticism of a simple bayesian linear regression model

When calling sess.run(Y\_post.mean(), feed\_dict=X\_tst) the result varies each time this is run. I suspect this has something to do with the variational parameter distributions not being set to return their mean. Is this correct and if one wanted exactly the same predictive mean would you have to specify the variational distributions to be the mean values?

i.e. when calling

> mean = sess.run(Y\_post.mean(), feed\_dict=X\_test)

Should the definintion of Y\_post be defined instead of

> Y\_post = ed.copy(Y, {W: qW, … } )

as

> Y\_post = ed.copy(Y, {W: qW.mean(), … } )

is this the correct way to do this? If not what is the recommended way to do this

---

<div class="post-metadata">

**Author:** ![dustin](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.edwardlib.org/dustin/32/134_2.png) [@dustin](https://discourse.edwardlib.org/u/dustin)\
**Post date:** [August 24, 2017, 7:30am UTC](https://discourse.edwardlib.org/t/prediction-or-criticism-of-the-model-mean/314/2 "2017-08-24T07:30:35Z")

</div>

Fetching `y_post` draws new parameters because any random variables it depends on in the computational graph are redrawn. The same happens if you try to fetch `x` in the program  
`theta = Beta(1.0, 1.0); x = Bernoulli(probs=theta, sample_shape=50)`.

> [@MushroomHunting](#):
>
> Should the definintion of Y\_post be defined instead of
> 
> Y\_post = ed.copy(Y, {W: qW, … } )
> 
> as
> 
> Y\_post = ed.copy(Y, {W: qW.mean(), … } )
> 
> is this the correct way to do this?

Consider what this means mathematically. The first line represents the posterior predictive,

```auto
p(xnew | x) = \int p(xnew | theta) p(theta | x) d\theta

```

The second line represents the likelihood with parameters given by the posterior mean, `p(xnew | theta = mean(p(theta | x)))`.

In general, to calculate something like the posterior predictive mean you should fetch `y_post` many times and average.

---

<div class="post-metadata">

**Author:** ![MushroomHunting](https://avatars.discourse-cdn.com/v4/letter/m/90ced4/32.png) [@MushroomHunting](https://discourse.edwardlib.org/u/MushroomHunting)\
**Post date:** [August 25, 2017, 4:37am UTC](https://discourse.edwardlib.org/t/prediction-or-criticism-of-the-model-mean/314/3 "2017-08-25T04:37:54Z")

</div>

Thanks Dustin!

To clarify, is there any difference in running Y\_post, Y\_post.mean() and multiple Y\_post.sample() and then averaging their results; with Y\_post.sample([num\_samples]) being the most efficient to obtain an approximate posterior mean?

---

<div class="post-metadata">

**Author:** ![dustin](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.edwardlib.org/dustin/32/134_2.png) [@dustin](https://discourse.edwardlib.org/u/dustin)\
**Post date:** [August 25, 2017, 11:24pm UTC](https://discourse.edwardlib.org/t/prediction-or-criticism-of-the-model-mean/314/4 "2017-08-25T23:24:17Z")

</div>

> [@MushroomHunting](#):
>
> To clarify, is there any difference in running Y\_post, Y\_post.mean() and multiple Y\_post.sample() and then averaging their results; with Y\_post.sample([num\_samples]) being the most efficient to obtain an approximate posterior mean?

Yes.

It’s worth working out what these mean:

- `np.mean([sess.run(y_post) for _ in range(50)])` fetches a posterior sample; then likelihood sample; then repeats 50 times and takes the mean.
- `sess.run(y_post.mean())` fetches a posterior sample, then takes the likelihood’s mean given the single posterior sample.
- `sess.run(y_post.sample([num_samples]))` fetches a posterior sample, then draws `num_samples` samples from the likelihood given the single posterior sample.

Only the first method is correct.

---

<div class="post-metadata">

**Author:** ![MushroomHunting](https://avatars.discourse-cdn.com/v4/letter/m/90ced4/32.png) [@MushroomHunting](https://discourse.edwardlib.org/u/MushroomHunting)\
**Post date:** [August 27, 2017, 9:07am UTC](https://discourse.edwardlib.org/t/prediction-or-criticism-of-the-model-mean/314/5 "2017-08-27T09:07:08Z")

</div>

cheers; this makes things clearer

---

<div class="post-metadata">

**Author:** ![Freezinger](https://avatars.discourse-cdn.com/v4/letter/f/54ee81/32.png) [@Freezinger](https://discourse.edwardlib.org/u/Freezinger)\
**Post date:** [July 11, 2018, 3:03pm UTC](https://discourse.edwardlib.org/t/prediction-or-criticism-of-the-model-mean/314/6 "2018-07-11T15:03:24Z")

</div>

Sorry to reawaken this old thread but I’ve been asking this question for a project I’ve been working on, too. In the case of variational inference in a linear model that consists strictly of independent normally distributed random variables, don’t all random variables converge in probability to their posterior means when sampling from y\_post? In this case, shouldn’t it be correct to evaluate the model using ed.copy() to replace all random variables (including y) with their posterior means and then generate a single prediction? This would just be for computing a point estimate of the model’s error/likelihood. For other types of model criticism the full posterior predictive would be needed.

The implementation would look like this:

```auto
# From the supervised learning tutorial
from edward.models import Normal

X = tf.placeholder(tf.float32, [N, D])
w = Normal(loc=tf.zeros(D), scale=tf.ones(D))
b = Normal(loc=tf.zeros(1), scale=tf.ones(1))
y = Normal(loc=ed.dot(X, w) + b, scale=tf.ones(N))

qw = Normal(loc=tf.get_variable("qw/loc", [D]),
            scale=tf.nn.softplus(tf.get_variable("qw/scale", [D])))
qb = Normal(loc=tf.get_variable("qb/loc", [1]),
            scale=tf.nn.softplus(tf.get_variable("qb/scale", [1])))

inference = ed.KLqp({w: qw, b: qb}, data={X: X_train, y: y_train})
inference.run(n_samples=5, n_iter=250)
y_post = ed.copy(y, {w: qw, b: qb})

# Current proposal, used only for point estimation of error/likelihood
y_MAP = ed.copy(y.mean(), {w: qw.mean(), b: qb.mean()}, scope='MAP')

```
