# Rule of thumb in choosing n\_samples

**URL:** <https://discourse.edwardlib.org/t/rule-of-thumb-in-choosing-n-samples/896>\
**Category:** General\
**Created:** [August 8, 2018, 4:45am UTC](https://discourse.edwardlib.org/t/rule-of-thumb-in-choosing-n-samples/896 "2018-08-08T04:45:23Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![haoyangz](https://avatars.discourse-cdn.com/v4/letter/h/3ab097/32.png) [@haoyangz](https://discourse.edwardlib.org/u/haoyangz)\
**Post date:** [August 8, 2018, 4:45am UTC](https://discourse.edwardlib.org/t/rule-of-thumb-in-choosing-n-samples/896/1 "2018-08-08T04:45:23Z")

</div>

If I understand correctly, `n_samples` in inference classes like `ed.KLqp` is the Monte Carlo sample number `S` in [here](http://edwardlib.org/tutorials/klqp). Is this correct?

What’s a good rule of thumb in choosing `n_sample`? The more the better within a time budget?

---

<div class="post-metadata">

**Author:** ![tanaka-hiroki1989](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.edwardlib.org/tanaka-hiroki1989/32/275_2.png) [@tanaka-hiroki1989](https://discourse.edwardlib.org/u/tanaka-hiroki1989)\
**Post date:** [August 8, 2018, 9:21am UTC](https://discourse.edwardlib.org/t/rule-of-thumb-in-choosing-n-samples/896/2 "2018-08-08T09:21:11Z")

</div>

That is correct.

However, I don’t know the good rule clearly.

In my experience, if the number of samples is large, the variance of the variational approximation distribution decreases, but convergence does not become faster.

---

<div class="post-metadata">

**Author:** ![dekabeyler](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.edwardlib.org/dekabeyler/32/381_2.png) [@dekabeyler](https://discourse.edwardlib.org/u/dekabeyler)\
**Post date:** [August 13, 2018, 6:34pm UTC](https://discourse.edwardlib.org/t/rule-of-thumb-in-choosing-n-samples/896/3 "2018-08-13T18:34:37Z")

</div>

In [https://github.com/blei-lab/edward/blob/master/notebooks/tensorboard.ipynb](https://github.com/blei-lab/edward/blob/master/notebooks/tensorboard.ipynb)

The authors state:

" With variational inference, we also include information such as the loss function and its decomposition into individual terms. This particular example shows that `n_samples=1` tends to have higher variance than `n_samples=5` but still converges to the same solution."

which supports your claim that larger number of MC samples decreases the variance.
