# Streamed batch training with manual update - scaling considerations

**URL:** <https://discourse.edwardlib.org/t/streamed-batch-training-with-manual-update-scaling-considerations/262>\
**Category:** General\
**Created:** [July 26, 2017, 11:24am UTC](https://discourse.edwardlib.org/t/streamed-batch-training-with-manual-update-scaling-considerations/262 "2017-07-26T11:24:07Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![MushroomHunting](https://avatars.discourse-cdn.com/v4/letter/m/90ced4/32.png) [@MushroomHunting](https://discourse.edwardlib.org/u/MushroomHunting)\
**Post date:** [July 26, 2017, 11:24am UTC](https://discourse.edwardlib.org/t/streamed-batch-training-with-manual-update-scaling-considerations/262/1 "2017-07-26T11:24:07Z")

</div>

Hi there

related:

> [@How to update model with new data?](https://discourse.edwardlib.org/t/how-to-update-model-with-new-data/216):
>
> Suppose that I initially train a Bayesian model with 5 training samples. As time passes by, I get more samples, and want to update the model with 2 samples a time for 20 times. Below is my code, I expect to see that the below result shoud be closer and closer to the true value as more samples are added and finally should be the same as I train the model with 45 samples together. But that’s not the case, the euclidean distance between the true and estimated values ocillate after the first 5 train…

> [@Iterative estimators ("bayes filters") in Edward?](https://discourse.edwardlib.org/t/iterative-estimators-bayes-filters-in-edward/104/2):
>
> To handle streaming data, the easiest place to get started might be with the [Bayesian linear regression](https://github.com/blei-lab/edward/blob/master/examples/bayesian_linear_regression_ppc.py) example. Instead of calling inference.run(), you can manually write the training loop: # placeholders on data, where we will pass in values during training X\_ph = tf.placeholder(tf.float32, [None, D]) y\_ph = tf.placeholder(tf.float32, [None]) inference = ed.KLqp({w: qw, b: qb}, data={X: X\_ph, y: y\_ph}) inference.initialize() tf.global\_variables\_initializer().run() for \_ in range(inference.…

My starting point is [Edward – Batch Training](http://edwardlib.org/tutorials/batch-training) in which the tutorial makes a note of setting a scale factor of

> N/M

Like the two related posts up the top i’m interested in running VI in a streaming context. Regarding the initialisation of ed.KLqp I have a few confusions

1. Is the scale parameter only relevant if one uses inference.run at first with N \> M training data points.

2. Is n\_samples only relevant if one uses an inference. run. call? I.e. if from the beginning if you have a large amount of data to train first on, and then the model is exposed to completely new data in a streaming situation. If not where does this fit in

3. In the case of having no data initially, and then acquiring data accumulating into buffers of size M, is it enough to simply have N=M giving a scale of 1 ?

cheers!

---

<div class="post-metadata">

**Author:** ![dustin](https://yyz1.discourse-cdn.com/flex035/user_avatar/discourse.edwardlib.org/dustin/32/134_2.png) [@dustin](https://discourse.edwardlib.org/u/dustin)\
**Post date:** [August 11, 2017, 4:03pm UTC](https://discourse.edwardlib.org/t/streamed-batch-training-with-manual-update-scaling-considerations/262/2 "2017-08-11T16:03:33Z")

</div>

> [@MushroomHunting](#):
>
> Is the scale parameter only relevant if one uses inference.run at first with N \> M training data points.

To some degree, yes. It’s used generally to scale any computation with respect to the random variables. For example, you might use it for masking.

> [@MushroomHunting](#):
>
> Is n\_samples only relevant if one uses an inference. run. call? I.e. if from the beginning if you have a large amount of data to train first on, and then the model is exposed to completely new data in a streaming situation. If not where does this fit in

`n_samples` is an algorithm hyperparameter in `ed.KLqp`, representing the number of samples to estimate the gradient of the loss function. It is always relevant whenever running the algorithm.

> [@MushroomHunting](#):
>
> In the case of having no data initially, and then acquiring data accumulating into buffers of size M, is it enough to simply have N=M giving a scale of 1 ?

That would only work if after each streaming batch, you re-set the prior distribution to be the inferred posterior from the batch. Otherwise, imagine you had a billion streaming points, with a batch size of 1; without scaling the likelihood by 1 million, the prior overwhelms the likelihood so the posterior will not differ much from the prior.

The discussion in [Iterative estimators ("bayes filters") in Edward? - #2 by dustin](https://discourse.edwardlib.org/t/iterative-estimators-bayes-filters-in-edward/104/2) provides most detail.

---

<div class="post-metadata">

**Author:** ![MushroomHunting](https://avatars.discourse-cdn.com/v4/letter/m/90ced4/32.png) [@MushroomHunting](https://discourse.edwardlib.org/u/MushroomHunting)\
**Post date:** [August 12, 2017, 12:31am UTC](https://discourse.edwardlib.org/t/streamed-batch-training-with-manual-update-scaling-considerations/262/3 "2017-08-12T00:31:18Z")

</div>

thanks Dustin, that clears things up
