<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>BAIR Blog &#8211; Robohub</title>
	<atom:link href="https://robohub.org/author/bairblog/feed/" rel="self" type="application/rss+xml" />
	<link>https://robohub.org</link>
	<description>Connecting the robotics community to the world</description>
	<lastBuildDate>Tue, 28 Apr 2026 07:40:00 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.3</generator>
	<item>
		<title>Gradient-based planning for world models at longer horizons</title>
		<link>https://robohub.org/gradient-based-planning-for-world-models-at-longer-horizons/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Tue, 28 Apr 2026 07:00:44 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://bair.berkeley.edu/blog/2026/04/20/grasp/</guid>

					<description><![CDATA[<!-- twitter -->














<div>
  <img src="https://bair.berkeley.edu/static/blog/grasp/ballnav_demo.gif" alt="BallNav demo">
  <img src="https://bair.berkeley.edu/static/blog/grasp/pusht_zoomout.gif" alt="Push-T demo">
</div>

<p><strong>GRASP</strong> is a new gradient-based planner for learned dynamics (a “world model”) that makes long-horizon planning practical by (1) lifting the trajectory into virtual states so optimization is parallel across time, (2) adding stochasticity directly to the state iterates for exploration, and (3) reshaping gradients so actions get clean signals while we avoid brittle “state-input” gradients through high-dimensional vision models.</p>

<!--more-->

<p>Large, learned world models are becoming increasingly capable. They can predict long sequences of future observations in high-dimensional visual spaces and generalize across tasks in ways that were difficult to imagine a few years ago. As these models scale, they start to look less like task-specific predictors and more like general-purpose simulators.</p>

<p>But having a powerful predictive model is not the same as being able to use it effectively for control/learning/planning. In practice, long-horizon planning with modern world models remains fragile: optimization becomes ill-conditioned, non-greedy structure creates bad local minima, and high-dimensional latent spaces introduce subtle failure modes.</p>

<p>In this blog post, I describe the problems that motivated this project and our approach to address them: why planning with modern world models can be surprisingly fragile, why long horizons are the real stress test, and what we changed to make gradient-based planning much more robust.</p>

<hr>

<blockquote>
  <p>This blog post discusses work done with Mike Rabbat, Aditi Krishnapriyan, Yann LeCun, and Amir Bar (* denotes equal advisorship), where we propose GRASP.</p>
</blockquote>

<hr>

<h2>What is a world model?</h2>

<p>These days, the term “world model” is quite overloaded, and depending on the context can either mean an explicit dynamics model or some implicit, reliable internal state that a generative model relies on (e.g. when an LLM generates chess moves, whether there is some internal representation of the board). We give our loose working definition below.</p>

<p>Suppose you take actions $a_t \in \mathcal{A}$ and observe states $s_t \in \mathcal{S}$ (images, latent vectors, proprioception). A <strong>world model</strong> is a learned model that, given the current state and a sequence of future actions, predicts what will happen next. Formally, it defines a predictive distribution on a sequence of observed states $s_{t-h:t}$ and current action $a_t$:</p>

\[P_\theta(s_{t+1} \mid s_{t-h:t},\; a_t)\]

<p>that approximates the environment’s true conditional $P(s_{t+1} \mid s_{t-h:t},\; a_t)$. For this blog post, we’ll assume a Markovian model $P(s_{t+1} \mid s_{t-h:t},\; a_t)$ for simplicity (all results here can be extended to the more general case), and when the model is deterministic it reduces to a map over states:</p>

\[s_{t+1} = F_\theta(s_t, a_t).\]

<p>In practice the state $s_t$ is often a learned latent representation (e.g., encoded from pixels), so the model operates in a (theoretically) compact, differentiable space. The key point is that a world model gives you a <em>differentiable simulator</em>; you can roll it forward under hypothetical action sequences and backpropagate through the predictions.</p>

<hr>

<h2>Planning: choosing actions by optimizing through the model</h2>

<p>Given a start $s_0$ and a goal $g$, the simplest planner chooses an action sequence $\mathbf{a}=(a_0,\dots,a_{T-1})$ by rolling out the model and minimizing terminal error:</p>

\[\min_{\mathbf{a}} \; \&#124; s_T(\mathbf{a}) - g \&#124;_2^2, \quad \text{where } s_T(\mathbf{a}) = \mathcal{F}_{\theta}^{T}(s_0,\mathbf{a}).\]

<p>Here we use $\mathcal{F}^T$ as shorthand for the full rollout through the world model (dependence on model parameters $\theta$ is implicit):</p>

\[\mathcal{F}_{\theta}^{T}(s_0, \mathbf{a}) = F_\theta(F_\theta(\cdots F_\theta(s_0, a_0), \cdots, a_{T-2}), a_{T-1}).\]

<p>In short horizons and low-dimensional systems, this can work reasonably well. But as horizons grow and models become larger and more expressive, its weaknesses become amplified.</p>

<p>So why doesn’t this just work at scale?</p>

<hr>

<h2>Why long-horizon planning is hard (even when everything is differentiable)</h2>

<p>There are two separate pain points for the more general world model, plus a third that is specific to learned, deep learning-based models.</p>

<h3>1) Long-horizon rollouts create deep, ill-conditioned computation graphs</h3>

<p>Those familiar with backprop through time (BPTT) may notice that we’re differentiating through a model applied to itself repeatedly, which will lead to the <strong>exploding/vanishing gradients</strong> problem. Namely, if we take derivatives (note we’re differentiating vector-valued functions, resulting in Jacobians that we denote with $D_x (\cdots)$) with respect to earlier actions (e.g. $a_0$):</p>

\[D_{a_0} \mathcal{F}_{\theta}^{T}(s_0, \mathbf{a}) = \Bigl(\prod_{t=1}^T D_s F_\theta(s_t, a_t)\Bigr) D_{a_0}F_\theta(s_0, a_0).\]

<p>We see that the Jacobian’s conditioning scales exponentially with time $T$:</p>

\[\sigma_{\text{max/min}}(D_{a_0}\mathcal{F}_{\theta}^{T}) \sim \sigma_{\text{max/min}}(D_s F_\theta)^{T-1},\]

<p>leading to exploding or vanishing gradients.</p>

<h3>2) The landscape is non-greedy and full of traps</h3>

<p>At short horizons, the greedy solution, where we move straight toward the goal at every step, is often good enough. If you only need to plan a few steps ahead, the optimal trajectory usually doesn’t deviate much from “head toward $g$” at each step.</p>

<p>As horizons grow, two things happen. First, longer tasks are more likely to require <em>non-greedy</em> behavior: going around a wall, repositioning before pushing, backing up to take a better path. And as horizons grow, more of these non-greedy steps are typically needed. Second, the optimization space itself scales with horizon: $\mathrm{dim}(\mathcal{A} \times \cdots \times \mathcal{A}) = T\mathrm{dim}(\mathcal{A})$, further expanding the space of local minima for the optimization problem.</p>

<figure>
  <img src="https://bair.berkeley.edu/static/blog/grasp/loss-landscape.jpg" alt="Loss landscape">
  <figcaption><em>Distance to goal along the optimal path is non-monotonic, and the resulting loss landscape can be rough.</em></figcaption>
</figure>

<hr>

<h2>A long-horizon fix: lifting the dynamics constraint</h2>

<p>Suppose we treat the dynamics constraint $s_{t+1} = F_{\theta}(s_t, a_t)$ as a soft constraint, and we instead optimize the following penalty function over both actions $(a_0,\ldots,a_{T-1})$ and states $(s_0,\ldots,s_T)$:</p>

\[\min_{\mathbf{s},\mathbf{a}} \mathcal{L}(\mathbf{s}, \mathbf{a}) = \sum_{t=0}^{T-1} \big\&#124;F_\theta(s_t,a_t) - s_{t+1}\big\&#124;_2^2,
\quad \text{with } s_0 \text{ fixed and } s_T=g.\]

<p>This is also sometimes called <em>collocation</em> in planning/robotics literature. Note the lifted formulation shares the same <em>global</em> minimizers as the original rollout objective (both are zero exactly when the trajectory is dynamically feasible). But the optimization landscapes are very different, and we get two immediate benefits:</p>

<ul>
  <li>Each world model evaluation $F_{\theta}(s_t,a_t)$ depends only on local variables, so all $T$ terms can be computed <em>in parallel across time</em>, resulting in a huge speed-up for longer horizons, and</li>
  <li>You no longer backpropagate through a single deep $T$-step composition to get a learning signal, since the previous product of Jacobians now splits into a sum, e.g.:</li>
</ul>

\[D_{a_0} \mathcal{L} = 2(F_\theta(s_0, a_0) - s_1).\]

<p>Being able to optimize states directly also helps with exploration, as we can temporarily navigate through unphysical domains to find the optimal plan:</p>
<figure>
  <img src="https://bair.berkeley.edu/static/blog/grasp/ballnav_demo.gif" alt="Collocation planning in BallNav">
  <figcaption><em>Collocation-based planning allows us to directly perturb states and explore midpoints more effectively.</em></figcaption>
</figure>

<p>However, lunch is never free. And indeed, especially for deep learning-based world models, there is a critical issue that makes the above optimization quite difficult in practice.</p>

<h2>An issue for deep learning-based world models: sensitivity of state-input gradients</h2>

<p>The <strong>tl;dr</strong> of this section is: directly optimizing states through a deep learning-based $F_{\theta}$ is incredibly brittle, à la <em>adversarial robustness</em>. Even if you train your world model in a lower-dimensional state space, the training process for the world model makes unseen state landscapes very sharp, whether it be an unseen state itself or simply a normal/orthogonal direction to the data manifold.</p>

<h3>Adversarial robustness and the “dimpled manifold” model</h3>

<p>Adversarial robustness originally looked at classification models $f_\theta : \mathbb{R}^{w\times h \times c} \to \mathbb{R}^K$, and showed that by following the gradient of a particular logit $\nabla f_\theta^k$ from a base image $x$ (not of class $k$), you did not have to move far along $x’ = x + \epsilon\nabla f_\theta^k$ to make $f_\theta$ classify $x’$ as $k$ (<a href="https://arxiv.org/abs/1312.6199" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Szegedy et al., 2014</a>; <a href="https://arxiv.org/abs/1412.6572" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Goodfellow et al., 2015</a>):</p>

<figure>
  <img src="https://bair.berkeley.edu/static/blog/grasp/adversarial_animated.gif" alt="Adversarial example">
  <figcaption><em>Depiction of the classic example from (Goodfellow et al., 2015).</em></figcaption>
</figure>

<p>Later work has painted a geometric picture for what’s going on: for data near a low-dimensional manifold $\mathcal{M}$, the training process controls behavior in tangential directions, but does not regularize behavior in orthogonal directions, thus leading to sensitive behavior (<a href="https://arxiv.org/pdf/1812.00740" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Stutz et al., 2019</a>). Another way stated: $f_\theta$ has a reasonable Lipschitz constant when considering only tangential directions to the data manifold $\mathcal{M}$, but can have very high Lipschitz constants in normal directions. In fact, it often benefits the model to be sharper in these normal directions, so it can fit more complicated functions more precisely.</p>

<figure>
  <img src="https://bair.berkeley.edu/static/blog/grasp/manifold_adversarial.gif" alt="Adversarial perturbations leave the data manifold">
</figure>

<p>As a result, such adversarial examples are incredibly common even for a single given model. Further, this is not just a computer vision phenomenon; adversarial examples also appear in LLMs (<a href="https://arxiv.org/abs/1908.07125" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wallace et al., 2019</a>) and in RL (<a href="https://arxiv.org/abs/1905.10615" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gleave et al., 2019</a>).</p>

<p>While there are methods to train for more adversarially robust models, there is a known trade-off between model performance and adversarial robustness (<a href="https://arxiv.org/pdf/1805.12152" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Tsipras et al., 2019</a>): especially in the presence of many weakly-correlated variables, the model <em>must</em> be sharper to achieve higher performance. Indeed, most modern training algorithms, whether in computer vision or LLMs, do not train adversarial robustness out. Thus, at least until deep learning sees a major regime change, <strong>this is a problem we’re stuck with</strong>.</p>

<h3>Why is adversarial robustness an issue for world model planning?</h3>

<p>Consider a single component of the dynamics loss we’re optimizing in the lifted state approach:</p>

\[\min_{s_t, a_t, s_{t+1}} \&#124;F_\theta(s_t, a_t) - s_{t+1}\&#124;_2^2\]

<p>Let’s further focus on just the base state:</p>

\[\min_{s_t} \&#124;F_\theta(s_t, a_t) - s_{t+1}\&#124;_2^2.\]

<p>Since world models are typically trained on state/action trajectories $(s_1, a_1, s_2, a_2, \ldots)$, the state-data manifold for $F_{\theta}$ has dimensionality bounded by the action space:</p>

\[\mathrm{dim}(\mathcal{M}_s) \le \mathrm{dim}(\mathcal{A}) + 1 + \mathrm{dim}(\mathcal{R}),\]

<p>where $\mathcal{R}$ is some optional space of augmentations (e.g. translations/rotations). Thus, we can typically expect $\mathrm{dim}(\mathcal{M}_s)$ to be much lower than $\mathrm{dim}(\mathcal{S})$, and thus: <strong>it is very easy to find adversarial examples that hack any state to any other desired state.</strong></p>

<p>As a result, the dynamics optimization</p>

\[\sum_{t=0}^{T-1} \big\&#124;F_\theta(s_t,a_t) - s_{t+1}\big\&#124;_2^2\]

<p>feels incredibly “sticky,” as the base points $s_t$ can easily trick $F_{\theta}$ into thinking it’s already made its local goal.<sup><a href="http://bair.berkeley.edu/blog/2026/04/20/grasp/#fn1" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">1</a></sup></p>

<figure>
  <img src="https://bair.berkeley.edu/static/blog/grasp/pusht_adversarial.gif" alt="Adversarial world model example">
</figure>

<hr>

<div>

  <p><strong>1.</strong> This adversarial robustness issue, while particularly bad for lifted-state approaches, is not unique to them. Even for serial optimization methods that optimize through the full rollout map $\mathcal{F}^T$, it is possible to get into unseen states, where it is very easy to have a normal component fed into the sensitive normal components of $D_s F_{\theta}$. The action Jacobian’s chain rule expansion is</p>

\[\Bigl(\prod_{t=1}^T D_s F_\theta(s_t, a_t)\Bigr) D_{a_0}F_\theta(s_0, a_0).\]

  <p>See what happens if any stage of the product has any component normal to the data manifold. <a href="http://bair.berkeley.edu/blog/2026/04/20/grasp/#ref1" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">↩</a></p>

</div>

<hr>

<h3>Our fix</h3>

<p>This is where our new planner GRASP comes in. The main observation: while $D_s F_{\theta}$ is untrustworthy and adversarial, the action space is usually low-dimensional and exhaustively trained, so $D_a F_{\theta}$ is actually reasonable to optimize through and doesn’t suffer from the adversarial robustness issue!</p>

<figure>
  <img src="https://bair.berkeley.edu/static/blog/grasp/network_diagram.jpg" alt="Network diagram showing high-dim state vs low-dim action">
  <figcaption><em>The action input is usually lower-dimensional and densely trained (the model has seen every action direction), so action gradients are much better behaved.</em></figcaption>
</figure>

<p>At its core, <strong>GRASP builds a first-order lifted state / collocation-based planner that is only dependent on action Jacobians through the world model.</strong> We thus exploit the differentiability of learned world models $F_{\theta}$, while not falling victim to the inherent sensitivity of the state Jacobians $D_s F_{\theta}$.</p>

<h2>GRASP: Gradient <strong>RelAxed</strong> <strong>S</strong>tochastic <strong>P</strong>lanner</h2>

<p>As noted before, we start with the collocation planning objective, where we lift the states and relax dynamics into a penalty:</p>

\[\min_{\mathbf{s},\mathbf{a}} \mathcal{L}(\mathbf{s}, \mathbf{a}) = \sum_{t=0}^{T-1} \big\&#124;F_\theta(s_t,a_t) - s_{t+1}\big\&#124;_2^2,
\quad \text{with } s_0 \text{ fixed and } s_T=g.\]

<p>We then make two key additions.</p>

<h2>Ingredient 1: Exploration by noising the <strong>state iterates</strong></h2>

<p>Even with a smoother objective, planning is nonconvex. We introduce exploration by injecting Gaussian noise into the <strong>virtual state updates</strong> during optimization.</p>

<p>A simple version:</p>

\[s_t \leftarrow s_t - \eta_s \nabla_{s_t}\mathcal{L} + \sigma_{\text{state}} \xi, \qquad \xi\sim\mathcal{N}(0,I).\]

<p>Actions are still updated by non-stochastic descent:</p>

\[a_t \leftarrow a_t - \eta_a \nabla_{a_t}\mathcal{L}.\]

<p>The state noise helps you “hop” between basins in the lifted space, while the actions remain guided by gradients. We found that specifically noising states here (as opposed to actions) finds a good balance of exploration and the ability to find sharper minima.<sup><a href="http://bair.berkeley.edu/blog/2026/04/20/grasp/#fn2" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">2</a></sup></p>

<hr>

<div>

  <p><strong>2.</strong> Because we only noise the states (and not the actions), the corresponding dynamics are not truly Langevin dynamics. <a href="http://bair.berkeley.edu/blog/2026/04/20/grasp/#ref2" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">↩</a></p>

</div>

<hr>

<h2>Ingredient 2: Reshape gradients: stop brittle state-input gradients, keep action gradients</h2>

<p>As discussed, the fragile pathway is the gradient that flows <em>into the state input</em> of the world model, <span>\(D_s F_{\theta}\)</span>. The most straightforward way to do this initially is to just stop state gradients into <span>\(F_{\theta}\)</span> directly:</p>

<ul>
  <li>Let $\bar{s}_t$ be the same value as $s_t$, but with gradients stopped.</li>
</ul>

<p>Define the <strong>stop-gradient dynamics loss</strong>:</p>

\[\mathcal{L}_{\text{dyn}}^{\text{sg}}(\mathbf{s},\mathbf{a})
= \sum_{t=0}^{T-1} \big\&#124;F_\theta(\bar{s}_t, a_t) - s_{t+1}\big\&#124;_2^2.\]

<p>This alone does not work. Notice now states only follow the previous state’s step, without anything forcing the base states to chase the next ones. As a result, there are trivial minima for just stopping at the origin, then only for the final action trying to get to the goal in one step.</p>

<h3>Dense goal shaping</h3>

<p>We can view the above issue as the goal’s signal being cut off entirely from previous states. One way to fix this is to simply add a dense goal term throughout prediction:</p>

\[\mathcal{L}_{\text{goal}}^{\text{sg}}(\mathbf{s},\mathbf{a})
= \sum_{t=0}^{T-1} \big\&#124;F_\theta(\bar{s}_t, a_t) - g\big\&#124;_2^2.\]

<p>In normal settings this would over-bias towards the greedy solution of straight chasing the goal, but this is balanced in our setting by the stop-gradient dynamics loss’s bias towards feasible dynamics. The final objective is then as follows:</p>

\[\mathcal{L}(\mathbf{s},\mathbf{a}) = \mathcal{L}_{\text{dyn}}^{\text{sg}}(\mathbf{s},\mathbf{a}) + \gamma \, \mathcal{L}_{\text{goal}}^{\text{sg}}(\mathbf{s},\mathbf{a}).\]

<p>The result is a planning optimization objective that does not have dependence on state gradients.</p>

<hr>

<h2>Periodic “sync”: briefly return to true rollout gradients</h2>

<p>The lifted stop-gradient objective is great for <strong>fast, guided exploration</strong>, but it’s still an approximation of the original serial rollout objective.</p>

<p>So every $K_{\text{sync}}$ iterations, GRASP does a short refinement phase:</p>

<ol>
  <li>Roll out from $s_0$ using current actions $\mathbf{a}$, and take a few small gradient steps on the original serial loss:</li>
</ol>

\[\mathbf{a} \leftarrow \mathbf{a} - \eta_{\text{sync}}\,\nabla_{\mathbf{a}}\,\&#124;s_T(\mathbf{a})-g\&#124;_2^2.\]

<p>The lifted-state optimization still provides the core of the optimization, while this refinement step adds some assistance to keep states and actions grounded towards real trajectories. This refinement step can of course be replaced with a serial planner of your choice (e.g. CEM); the core idea is to still get some of the benefit of the full-path synchronization of serial planners, while still mostly using the benefits of the lifted-state planning.</p>

<hr>

<h2>How GRASP addresses long-range planning</h2>

<p>Collocation-based planners offer a natural fix for long-horizon planning, but this optimization is quite difficult through modern world models due to adversarial robustness issues. <em>GRASP proposes a simple solution for a smoother collocation-based planner, alongside stable stochasticity for exploration</em>. As a result, longer-horizon planning ends up not only succeeding more, but also finding such successes faster:</p>

<figure>
  <img src="https://bair.berkeley.edu/static/blog/grasp/pusht_zoomout.gif" alt="Push-T planning demo">
  <figcaption><em>Push-T demo: longer-horizon planning with GRASP.</em></figcaption>
</figure>

<div class="grasp-results-table">

  <table>
    <thead>
      <tr>
        <th>Horizon</th>
        <th>CEM</th>
        <th>GD</th>
        <th>LatCo</th>
        <th><strong>GRASP</strong></th>
      </tr>
    </thead>
    <tbody>
      <tr>
        <td>H=40</td>
        <td><strong>61.4%</strong> / 35.3s</td>
        <td>51.0% / 18.0s</td>
        <td>15.0% / 598.0s</td>
        <td>59.0% / <strong>8.5s</strong></td>
      </tr>
      <tr>
        <td>H=50</td>
        <td>30.2% / 96.2s</td>
        <td>37.6% / 76.3s</td>
        <td>4.2% / 1114.7s</td>
        <td><strong>43.4%</strong> / <strong>15.2s</strong></td>
      </tr>
      <tr>
        <td>H=60</td>
        <td>7.2% / 83.1s</td>
        <td>16.4% / 146.5s</td>
        <td>2.0% / 231.5s</td>
        <td><strong>26.2%</strong> / <strong>49.1s</strong></td>
      </tr>
      <tr>
        <td>H=70</td>
        <td>7.8% / 156.1s</td>
        <td>12.0% / 103.1s</td>
        <td>0.0% / —</td>
        <td><strong>16.0%</strong> / <strong>79.9s</strong></td>
      </tr>
      <tr>
        <td>H=80</td>
        <td>2.8% / 132.2s</td>
        <td>6.4% / 161.3s</td>
        <td>0.0% / —</td>
        <td><strong>10.4%</strong> / <strong>58.9s</strong></td>
      </tr>
    </tbody>
  </table>

</div>

<p><em>Push-T results. Success rate (%) / median time to success. Bold = best in row. Note the median success time will bias higher with higher success rate; GRASP manages to be faster despite higher success rate.</em></p>

<hr>

<h2>What’s next?</h2>

<p>There is still plenty of work to be done for modern world model planners. We want to exploit the gradient structure of learned world models, and collocation (lifted-state optimization) is a natural approach for long-horizon planning, but it’s crucial to understand typical gradient structure here: smooth and informative action gradients and brittle state gradients. We view GRASP as an initial iteration for such planners.</p>

<p>Extension to diffusion-based world models (deeper latent timesteps can be viewed as smoothed versions of the world model itself), more sophisticated optimizers and noising strategies, and integrating GRASP into either a closed-loop system or RL policy learning for adaptive long-horizon planning are all natural and interesting next steps.</p>

<p>I do genuinely think it’s an exciting time to be working on world model planners. It’s a funny sweet spot where the background literature (planning and control overall) is incredibly mature and well-developed, but the current setting (pure planning optimization over modern, large-scale world models) is still heavily underexplored. But, once we figure out all the right ideas, world model planners will likely become as commonplace as RL.</p>

<hr>

<p>For more details, read the <a href="https://arxiv.org/pdf/2602.00475" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">full paper</a> or visit the <a href="https://www.michaelpsenka.io/grasp/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">project website</a>.</p>

<hr>

<h2>Citation</h2>

<div class="language-bibtex highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="nc">@article</span><span class="p">{</span><span class="nl">psenka2026grasp</span><span class="p">,</span>
  <span class="na">title</span><span class="p">=</span><span class="s">{Parallel Stochastic Gradient-Based Planning for World Models}</span><span class="p">,</span>
  <span class="na">author</span><span class="p">=</span><span class="s">{Michael Psenka and Michael Rabbat and Aditi Krishnapriyan and Yann LeCun and Amir Bar}</span><span class="p">,</span>
  <span class="na">year</span><span class="p">=</span><span class="s">{2026}</span><span class="p">,</span>
  <span class="na">eprint</span><span class="p">=</span><span class="s">{2602.00475}</span><span class="p">,</span>
  <span class="na">archivePrefix</span><span class="p">=</span><span class="s">{arXiv}</span><span class="p">,</span>
  <span class="na">primaryClass</span><span class="p">=</span><span class="s">{cs.LG}</span><span class="p">,</span>
  <span class="na">url</span><span class="p">=</span><span class="s">{https://arxiv.org/abs/2602.00475}</span>
<span class="p">}</span>
</code></pre></div></div>]]></description>
										<content:encoded><![CDATA[<div style="display: flex; flex-direction: column; align-items: center; gap: 1em; margin-bottom: 1.5em;">
  <img decoding="async" src="https://bair.berkeley.edu/static/blog/grasp/ballnav_demo.gif" alt="BallNav demo" style="max-width: 60%;" /><br />
  <img decoding="async" src="https://bair.berkeley.edu/static/blog/grasp/pusht_zoomout.gif" alt="Push-T demo" style="max-width: 90%;" />
</div>
<p><strong>By <a href="https://michael-psenka.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Michael Psenka</a>, <a href="https://ai.meta.com/people/1148536089838617/michael-rabbat/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Mike Rabbat</a>, <a href="https://a1k12.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Aditi Krishnapriyan</a>, <a href="https://yann.lecun.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Yann LeCun</a>, <a href="https://amirbar.net/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Amir Bar</a></strong></p>
<p><strong>GRASP</strong> is a new gradient-based planner for learned dynamics (a “world model”) that makes long-horizon planning practical by (1) lifting the trajectory into virtual states so optimization is parallel across time, (2) adding stochasticity directly to the state iterates for exploration, and (3) reshaping gradients so actions get clean signals while we avoid brittle “state-input” gradients through high-dimensional vision models.</p>
<p>Large, learned world models are becoming increasingly capable. They can predict long sequences of future observations in high-dimensional visual spaces and generalize across tasks in ways that were difficult to imagine a few years ago. As these models scale, they start to look less like task-specific predictors and more like general-purpose simulators.</p>
<p>But having a powerful predictive model is not the same as being able to use it effectively for control/learning/planning. In practice, long-horizon planning with modern world models remains fragile: optimization becomes ill-conditioned, non-greedy structure creates bad local minima, and high-dimensional latent spaces introduce subtle failure modes.</p>
<p>In this blog post, I describe the problems that motivated this project and our approach to address them: why planning with modern world models can be surprisingly fragile, why long horizons are the real stress test, and what we changed to make gradient-based planning much more robust.</p>
<hr />
<blockquote>
<p>This blog post discusses work done with Mike Rabbat, Aditi Krishnapriyan, Yann LeCun, and Amir Bar (* denotes equal advisorship), where we propose GRASP.</p>
</blockquote>
<hr />
<h2 id="what-is-a-world-model">What is a world model?</h2>
<p>These days, the term “world model” is quite overloaded, and depending on the context can either mean an explicit dynamics model or some implicit, reliable internal state that a generative model relies on (e.g. when an LLM generates chess moves, whether there is some internal representation of the board). We give our loose working definition below.</p>
<p>Suppose you take actions <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-a3c54d08fc60a73c5415736b15212106_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#97;&#95;&#116;&#32;&#92;&#105;&#110;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#65;&#125;" title="Rendered by QuickLaTeX.com" height="16" width="52" style="vertical-align: -3px;"/> and observe states <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-8f69e990aa1d6ebdae714e913745d6fb_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#115;&#95;&#116;&#32;&#92;&#105;&#110;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#83;&#125;" title="Rendered by QuickLaTeX.com" height="15" width="48" style="vertical-align: -3px;"/> (images, latent vectors, proprioception). A <strong>world model</strong> is a learned model that, given the current state and a sequence of future actions, predicts what will happen next. Formally, it defines a predictive distribution on a sequence of observed states <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-88b4b5b7d2b439238442a7aa5e01cb86_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#115;&#95;&#123;&#116;&#45;&#104;&#58;&#116;&#125;" title="Rendered by QuickLaTeX.com" height="11" width="41" style="vertical-align: -3px;"/> and current action <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-b5f8b679d5b1ba8241ec391b34717ef0_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#97;&#95;&#116;" title="Rendered by QuickLaTeX.com" height="11" width="14" style="vertical-align: -3px;"/>:</p>
<p class="ql-center-displayed-equation" style="line-height: 19px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-95106aed300ef1fe3763ed3e7a4c35a2_l3.png" height="19" width="148" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#80;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#123;&#116;&#43;&#49;&#125;&#32;&#92;&#109;&#105;&#100;&#32;&#115;&#95;&#123;&#116;&#45;&#104;&#58;&#116;&#125;&#44;&#92;&#59;&#32;&#97;&#95;&#116;&#41;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>that approximates the environment’s true conditional <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-68d611727a003284884c7b755480587a_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#80;&#40;&#115;&#95;&#123;&#116;&#43;&#49;&#125;&#32;&#92;&#109;&#105;&#100;&#32;&#115;&#95;&#123;&#116;&#45;&#104;&#58;&#116;&#125;&#44;&#92;&#59;&#32;&#97;&#95;&#116;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="143" style="vertical-align: -5px;"/>. For this blog post, we’ll assume a Markovian model <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-68d611727a003284884c7b755480587a_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#80;&#40;&#115;&#95;&#123;&#116;&#43;&#49;&#125;&#32;&#92;&#109;&#105;&#100;&#32;&#115;&#95;&#123;&#116;&#45;&#104;&#58;&#116;&#125;&#44;&#92;&#59;&#32;&#97;&#95;&#116;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="143" style="vertical-align: -5px;"/> for simplicity (all results here can be extended to the more general case), and when the model is deterministic it reduces to a map over states:</p>
<p class="ql-center-displayed-equation" style="line-height: 19px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-d6e5828483a4880c1cfe7bad44d78973_l3.png" height="19" width="129" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#115;&#95;&#123;&#116;&#43;&#49;&#125;&#32;&#61;&#32;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#116;&#44;&#32;&#97;&#95;&#116;&#41;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>In practice the state <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-e86795deb37ff5f5055e741b17eb25d7_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#115;&#95;&#116;" title="Rendered by QuickLaTeX.com" height="11" width="13" style="vertical-align: -3px;"/> is often a learned latent representation (e.g., encoded from pixels), so the model operates in a (theoretically) compact, differentiable space. The key point is that a world model gives you a <em>differentiable simulator</em>; you can roll it forward under hypothetical action sequences and backpropagate through the predictions.</p>
<hr />
<h2 id="planning-choosing-actions-by-optimizing-through-the-model">Planning: choosing actions by optimizing through the model</h2>
<p>Given a start <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-7730c5a8c9bd4b8291e5231082f6b9c6_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#115;&#95;&#48;" title="Rendered by QuickLaTeX.com" height="11" width="15" style="vertical-align: -3px;"/> and a goal <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-d208fd391fa57c168dc0f151de829fee_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#103;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: -4px;"/>, the simplest planner chooses an action sequence <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-17f8a318ee1fb3a1f9da386869647837_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#61;&#40;&#97;&#95;&#48;&#44;&#92;&#100;&#111;&#116;&#115;&#44;&#97;&#95;&#123;&#84;&#45;&#49;&#125;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="141" style="vertical-align: -5px;"/> by rolling out the model and minimizing terminal error:</p>
<p class="ql-center-displayed-equation" style="line-height: 28px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-6dce57a108f3a4db376233c1e545dfd6_l3.png" height="28" width="357" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#109;&#105;&#110;&#95;&#123;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#125;&#32;&#92;&#59;&#32;&#92;&#124;&#32;&#115;&#95;&#84;&#40;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41;&#32;&#45;&#32;&#103;&#32;&#92;&#124;&#95;&#50;&#94;&#50;&#44;&#32;&#92;&#113;&#117;&#97;&#100;&#32;&#92;&#116;&#101;&#120;&#116;&#123;&#119;&#104;&#101;&#114;&#101;&#32;&#125;&#32;&#115;&#95;&#84;&#40;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41;&#32;&#61;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#70;&#125;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;&#94;&#123;&#84;&#125;&#40;&#115;&#95;&#48;&#44;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>Here we use <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-6b9827b4af48fd65e7e83838bfd859f6_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#70;&#125;&#94;&#84;" title="Rendered by QuickLaTeX.com" height="16" width="25" style="vertical-align: -1px;"/> as shorthand for the full rollout through the world model (dependence on model parameters <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-356a08e839ab6974a16448e16e56745d_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#116;&#104;&#101;&#116;&#97;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: 0px;"/> is implicit):</p>
<p class="ql-center-displayed-equation" style="line-height: 22px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-a46e0710d3438d8afddd4362bc5e468a_l3.png" height="22" width="389" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#70;&#125;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;&#94;&#123;&#84;&#125;&#40;&#115;&#95;&#48;&#44;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41;&#32;&#61;&#32;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#92;&#99;&#100;&#111;&#116;&#115;&#32;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#48;&#44;&#32;&#97;&#95;&#48;&#41;&#44;&#32;&#92;&#99;&#100;&#111;&#116;&#115;&#44;&#32;&#97;&#95;&#123;&#84;&#45;&#50;&#125;&#41;&#44;&#32;&#97;&#95;&#123;&#84;&#45;&#49;&#125;&#41;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>In short horizons and low-dimensional systems, this can work reasonably well. But as horizons grow and models become larger and more expressive, its weaknesses become amplified.</p>
<p>So why doesn’t this just work at scale?</p>
<hr />
<h2 id="why-long-horizon-planning-is-hard-even-when-everything-is-differentiable">Why long-horizon planning is hard (even when everything is differentiable)</h2>
<p>There are two separate pain points for the more general world model, plus a third that is specific to learned, deep learning-based models.</p>
<h3 id="1-long-horizon-rollouts-create-deep-ill-conditioned-computation-graphs">1) Long-horizon rollouts create deep, ill-conditioned computation graphs</h3>
<p>Those familiar with backprop through time (BPTT) may notice that we’re differentiating through a model applied to itself repeatedly, which will lead to the <strong>exploding/vanishing gradients</strong> problem. Namely, if we take derivatives (note we’re differentiating vector-valued functions, resulting in Jacobians that we denote with <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-b87824c6689dbd809ac8ebfdf787d8df_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#68;&#95;&#120;&#32;&#40;&#92;&#99;&#100;&#111;&#116;&#115;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="60" style="vertical-align: -5px;"/>) with respect to earlier actions (e.g. <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-d093ecd207b3f9d9b1c9cf0692b77d01_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#97;&#95;&#48;" title="Rendered by QuickLaTeX.com" height="11" width="16" style="vertical-align: -3px;"/>):</p>
<p class="ql-center-displayed-equation" style="line-height: 52px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-006a676ea89bf4f534c4e46f3822e638_l3.png" height="52" width="372" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#68;&#95;&#123;&#97;&#95;&#48;&#125;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#70;&#125;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;&#94;&#123;&#84;&#125;&#40;&#115;&#95;&#48;&#44;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41;&#32;&#61;&#32;&#92;&#66;&#105;&#103;&#108;&#40;&#92;&#112;&#114;&#111;&#100;&#95;&#123;&#116;&#61;&#49;&#125;&#94;&#84;&#32;&#68;&#95;&#115;&#32;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#116;&#44;&#32;&#97;&#95;&#116;&#41;&#92;&#66;&#105;&#103;&#114;&#41;&#32;&#68;&#95;&#123;&#97;&#95;&#48;&#125;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#48;&#44;&#32;&#97;&#95;&#48;&#41;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>We see that the Jacobian’s conditioning scales exponentially with time <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-f9ed275b0bf1633b7ee83b78fcc28273_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#84;" title="Rendered by QuickLaTeX.com" height="12" width="13" style="vertical-align: 0px;"/>:</p>
<p class="ql-center-displayed-equation" style="line-height: 24px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-7954df6601577938b25130705f47a19b_l3.png" height="24" width="312" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#115;&#105;&#103;&#109;&#97;&#95;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#109;&#97;&#120;&#47;&#109;&#105;&#110;&#125;&#125;&#40;&#68;&#95;&#123;&#97;&#95;&#48;&#125;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#70;&#125;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;&#94;&#123;&#84;&#125;&#41;&#32;&#92;&#115;&#105;&#109;&#32;&#92;&#115;&#105;&#103;&#109;&#97;&#95;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#109;&#97;&#120;&#47;&#109;&#105;&#110;&#125;&#125;&#40;&#68;&#95;&#115;&#32;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#41;&#94;&#123;&#84;&#45;&#49;&#125;&#44;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>leading to exploding or vanishing gradients.</p>
<h3 id="2-the-landscape-is-non-greedy-and-full-of-traps">2) The landscape is non-greedy and full of traps</h3>
<p>At short horizons, the greedy solution, where we move straight toward the goal at every step, is often good enough. If you only need to plan a few steps ahead, the optimal trajectory usually doesn’t deviate much from “head toward <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-d208fd391fa57c168dc0f151de829fee_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#103;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: -4px;"/>” at each step.</p>
<p>As horizons grow, two things happen. First, longer tasks are more likely to require <em>non-greedy</em> behavior: going around a wall, repositioning before pushing, backing up to take a better path. And as horizons grow, more of these non-greedy steps are typically needed. Second, the optimization space itself scales with horizon: <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-4bbae29ae0131c6238e884beb47ed8bf_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#114;&#109;&#123;&#100;&#105;&#109;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#65;&#125;&#32;&#92;&#116;&#105;&#109;&#101;&#115;&#32;&#92;&#99;&#100;&#111;&#116;&#115;&#32;&#92;&#116;&#105;&#109;&#101;&#115;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#65;&#125;&#41;&#32;&#61;&#32;&#84;&#92;&#109;&#97;&#116;&#104;&#114;&#109;&#123;&#100;&#105;&#109;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#65;&#125;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="229" style="vertical-align: -5px;"/>, further expanding the space of local minima for the optimization problem.</p>
<p><figure style="text-align: center;">
  <img decoding="async" src="https://bair.berkeley.edu/static/blog/grasp/loss-landscape.jpg" alt="Loss landscape" style="max-width: 80%;" /><figcaption><em>Distance to goal along the optimal path is non-monotonic, and the resulting loss landscape can be rough.</em></figcaption></figure>
</p>
<hr />
<h2 id="a-long-horizon-fix-lifting-the-dynamics-constraint">A long-horizon fix: lifting the dynamics constraint</h2>
<p>Suppose we treat the dynamics constraint <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-02d5b78079d86b85e769749919de58d4_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#115;&#95;&#123;&#116;&#43;&#49;&#125;&#32;&#61;&#32;&#70;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;&#40;&#115;&#95;&#116;&#44;&#32;&#97;&#95;&#116;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="124" style="vertical-align: -5px;"/> as a soft constraint, and we instead optimize the following penalty function over both actions <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-fe019867f244e7b47fe6ab35584574ab_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#40;&#97;&#95;&#48;&#44;&#92;&#108;&#100;&#111;&#116;&#115;&#44;&#97;&#95;&#123;&#84;&#45;&#49;&#125;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="106" style="vertical-align: -5px;"/> and states <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-1f06604430c41599fcbbbbb0372853f4_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#40;&#115;&#95;&#48;&#44;&#92;&#108;&#100;&#111;&#116;&#115;&#44;&#115;&#95;&#84;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="86" style="vertical-align: -5px;"/>:</p>
<p class="ql-center-displayed-equation" style="line-height: 52px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-863f11a3d6371cdc4342477c54c6f78f_l3.png" height="52" width="510" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#109;&#105;&#110;&#95;&#123;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#115;&#125;&#44;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#125;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#115;&#125;&#44;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41;&#32;&#61;&#32;&#92;&#115;&#117;&#109;&#95;&#123;&#116;&#61;&#48;&#125;&#94;&#123;&#84;&#45;&#49;&#125;&#32;&#92;&#98;&#105;&#103;&#92;&#124;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#116;&#44;&#97;&#95;&#116;&#41;&#32;&#45;&#32;&#115;&#95;&#123;&#116;&#43;&#49;&#125;&#92;&#98;&#105;&#103;&#92;&#124;&#95;&#50;&#94;&#50;&#44; &#92;&#113;&#117;&#97;&#100;&#32;&#92;&#116;&#101;&#120;&#116;&#123;&#119;&#105;&#116;&#104;&#32;&#125;&#32;&#115;&#95;&#48;&#32;&#92;&#116;&#101;&#120;&#116;&#123;&#32;&#102;&#105;&#120;&#101;&#100;&#32;&#97;&#110;&#100;&#32;&#125;&#32;&#115;&#95;&#84;&#61;&#103;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>This is also sometimes called <em>collocation</em> in planning/robotics literature. Note the lifted formulation shares the same <em>global</em> minimizers as the original rollout objective (both are zero exactly when the trajectory is dynamically feasible). But the optimization landscapes are very different, and we get two immediate benefits:</p>
<ul>
<li>Each world model evaluation <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-7643455d4474b4d26cbf516bd6dbd863_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#70;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;&#40;&#115;&#95;&#116;&#44;&#97;&#95;&#116;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="69" style="vertical-align: -5px;"/> depends only on local variables, so all <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-f9ed275b0bf1633b7ee83b78fcc28273_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#84;" title="Rendered by QuickLaTeX.com" height="12" width="13" style="vertical-align: 0px;"/> terms can be computed <em>in parallel across time</em>, resulting in a huge speed-up for longer horizons, and</li>
<li>You no longer backpropagate through a single deep <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-f9ed275b0bf1633b7ee83b78fcc28273_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#84;" title="Rendered by QuickLaTeX.com" height="12" width="13" style="vertical-align: 0px;"/>-step composition to get a learning signal, since the previous product of Jacobians now splits into a sum, e.g.:</li>
</ul>
<p class="ql-center-displayed-equation" style="line-height: 19px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-0ef7c03c999d91f9a56d87d2c2348184_l3.png" height="19" width="203" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#68;&#95;&#123;&#97;&#95;&#48;&#125;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#32;&#61;&#32;&#50;&#40;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#48;&#44;&#32;&#97;&#95;&#48;&#41;&#32;&#45;&#32;&#115;&#95;&#49;&#41;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>Being able to optimize states directly also helps with exploration, as we can temporarily navigate through unphysical domains to find the optimal plan:</p>
<p><figure style="text-align: center;">
  <img decoding="async" src="https://bair.berkeley.edu/static/blog/grasp/ballnav_demo.gif" alt="Collocation planning in BallNav" style="max-width: 60%;" /><figcaption><em>Collocation-based planning allows us to directly perturb states and explore midpoints more effectively.</em></figcaption></figure>
</p>
<p>However, lunch is never free. And indeed, especially for deep learning-based world models, there is a critical issue that makes the above optimization quite difficult in practice.</p>
<h2 id="an-issue-for-deep-learning-based-world-models-sensitivity-of-state-input-gradients">An issue for deep learning-based world models: sensitivity of state-input gradients</h2>
<p>The <strong>tl;dr</strong> of this section is: directly optimizing states through a deep learning-based <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-c502880341922d69202070782fbc9ab3_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#70;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;" title="Rendered by QuickLaTeX.com" height="15" width="18" style="vertical-align: -3px;"/> is incredibly brittle, à la <em>adversarial robustness</em>. Even if you train your world model in a lower-dimensional state space, the training process for the world model makes unseen state landscapes very sharp, whether it be an unseen state itself or simply a normal/orthogonal direction to the data manifold.</p>
<h3 id="adversarial-robustness-and-the-dimpled-manifold-model">Adversarial robustness and the “dimpled manifold” model</h3>
<p>Adversarial robustness originally looked at classification models <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-9cd4886e5f6ad5b59898d6a66094ee7b_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#102;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#32;&#58;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#82;&#125;&#94;&#123;&#119;&#92;&#116;&#105;&#109;&#101;&#115;&#32;&#104;&#32;&#92;&#116;&#105;&#109;&#101;&#115;&#32;&#99;&#125;&#32;&#92;&#116;&#111;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#82;&#125;&#94;&#75;" title="Rendered by QuickLaTeX.com" height="19" width="144" style="vertical-align: -4px;"/>, and showed that by following the gradient of a particular logit <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-fd14444ec7e338a8be9f26d1e8fc6e3c_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#110;&#97;&#98;&#108;&#97;&#32;&#102;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#94;&#107;" title="Rendered by QuickLaTeX.com" height="20" width="32" style="vertical-align: -5px;"/> from a base image <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ede05c264bba0eda080918aaa09c4658_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#120;" title="Rendered by QuickLaTeX.com" height="8" width="10" style="vertical-align: 0px;"/> (not of class <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3422b6bb5c160593658b7c39425d9880_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#107;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: 0px;"/>), you did not have to move far along <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-135dec93ed3ba20d5a0e6949123f416a_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#120;&#39;&#32;&#61;&#32;&#120;&#32;&#43;&#32;&#92;&#101;&#112;&#115;&#105;&#108;&#111;&#110;&#92;&#110;&#97;&#98;&#108;&#97;&#32;&#102;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#94;&#107;" title="Rendered by QuickLaTeX.com" height="20" width="110" style="vertical-align: -5px;"/> to make <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-6a20c3b1d68ef800563a48d91b7289d5_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#102;&#95;&#92;&#116;&#104;&#101;&#116;&#97;" title="Rendered by QuickLaTeX.com" height="16" width="16" style="vertical-align: -4px;"/> classify <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-b5e2f0a82567597c38101d0774b3fa68_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#120;&#39;" title="Rendered by QuickLaTeX.com" height="14" width="14" style="vertical-align: 0px;"/> as <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3422b6bb5c160593658b7c39425d9880_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#107;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: 0px;"/> (<a href="https://arxiv.org/abs/1312.6199" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Szegedy et al., 2014</a>; <a href="https://arxiv.org/abs/1412.6572" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Goodfellow et al., 2015</a>):</p>
<p><figure style="text-align: center;">
  <img decoding="async" src="https://bair.berkeley.edu/static/blog/grasp/adversarial_animated.gif" alt="Adversarial example" style="max-width: 70%;" /><figcaption><em>Depiction of the classic example from (Goodfellow et al., 2015).</em></figcaption></figure>
</p>
<p>Later work has painted a geometric picture for what’s going on: for data near a low-dimensional manifold <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-8b54f3b7741c19693e1e9d187786f082_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#77;&#125;" title="Rendered by QuickLaTeX.com" height="13" width="20" style="vertical-align: -1px;"/>, the training process controls behavior in tangential directions, but does not regularize behavior in orthogonal directions, thus leading to sensitive behavior (<a href="https://arxiv.org/pdf/1812.00740" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Stutz et al., 2019</a>). Another way stated: <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-6a20c3b1d68ef800563a48d91b7289d5_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#102;&#95;&#92;&#116;&#104;&#101;&#116;&#97;" title="Rendered by QuickLaTeX.com" height="16" width="16" style="vertical-align: -4px;"/> has a reasonable Lipschitz constant when considering only tangential directions to the data manifold <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-8b54f3b7741c19693e1e9d187786f082_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#77;&#125;" title="Rendered by QuickLaTeX.com" height="13" width="20" style="vertical-align: -1px;"/>, but can have very high Lipschitz constants in normal directions. In fact, it often benefits the model to be sharper in these normal directions, so it can fit more complicated functions more precisely.</p>
<p><figure style="text-align: center;">
  <img decoding="async" src="https://bair.berkeley.edu/static/blog/grasp/manifold_adversarial.gif" alt="Adversarial perturbations leave the data manifold" style="max-width: 70%;" /><br />
</figure>
</p>
<p>As a result, such adversarial examples are incredibly common even for a single given model. Further, this is not just a computer vision phenomenon; adversarial examples also appear in LLMs (<a href="https://arxiv.org/abs/1908.07125" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wallace et al., 2019</a>) and in RL (<a href="https://arxiv.org/abs/1905.10615" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gleave et al., 2019</a>).</p>
<p>While there are methods to train for more adversarially robust models, there is a known trade-off between model performance and adversarial robustness (<a href="https://arxiv.org/pdf/1805.12152" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Tsipras et al., 2019</a>): especially in the presence of many weakly-correlated variables, the model <em>must</em> be sharper to achieve higher performance. Indeed, most modern training algorithms, whether in computer vision or LLMs, do not train adversarial robustness out. Thus, at least until deep learning sees a major regime change, <strong>this is a problem we’re stuck with</strong>.</p>
<h3 id="why-is-adversarial-robustness-an-issue-for-world-model-planning">Why is adversarial robustness an issue for world model planning?</h3>
<p>Consider a single component of the dynamics loss we’re optimizing in the lifted state approach:</p>
<p class="ql-center-displayed-equation" style="line-height: 31px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-c0125bf4e6fc638ffbe447c238df12b1_l3.png" height="31" width="210" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#109;&#105;&#110;&#95;&#123;&#115;&#95;&#116;&#44;&#32;&#97;&#95;&#116;&#44;&#32;&#115;&#95;&#123;&#116;&#43;&#49;&#125;&#125;&#32;&#92;&#124;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#116;&#44;&#32;&#97;&#95;&#116;&#41;&#32;&#45;&#32;&#115;&#95;&#123;&#116;&#43;&#49;&#125;&#92;&#124;&#95;&#50;&#94;&#50;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>Let’s further focus on just the base state:</p>
<p class="ql-center-displayed-equation" style="line-height: 29px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-0dc167be31acb8f939fc7634fc9b3296_l3.png" height="29" width="185" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#109;&#105;&#110;&#95;&#123;&#115;&#95;&#116;&#125;&#32;&#92;&#124;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#116;&#44;&#32;&#97;&#95;&#116;&#41;&#32;&#45;&#32;&#115;&#95;&#123;&#116;&#43;&#49;&#125;&#92;&#124;&#95;&#50;&#94;&#50;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>Since world models are typically trained on state/action trajectories <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-56946ab92a181c0c5f0d12c3eccf56e9_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#40;&#115;&#95;&#49;&#44;&#32;&#97;&#95;&#49;&#44;&#32;&#115;&#95;&#50;&#44;&#32;&#97;&#95;&#50;&#44;&#32;&#92;&#108;&#100;&#111;&#116;&#115;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="130" style="vertical-align: -5px;"/>, the state-data manifold for <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-c502880341922d69202070782fbc9ab3_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#70;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;" title="Rendered by QuickLaTeX.com" height="15" width="18" style="vertical-align: -3px;"/> has dimensionality bounded by the action space:</p>
<p class="ql-center-displayed-equation" style="line-height: 19px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-fc6404df2a87ce26da2fffd7519bba15_l3.png" height="19" width="267" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#109;&#97;&#116;&#104;&#114;&#109;&#123;&#100;&#105;&#109;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#77;&#125;&#95;&#115;&#41;&#32;&#92;&#108;&#101;&#32;&#92;&#109;&#97;&#116;&#104;&#114;&#109;&#123;&#100;&#105;&#109;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#65;&#125;&#41;&#32;&#43;&#32;&#49;&#32;&#43;&#32;&#92;&#109;&#97;&#116;&#104;&#114;&#109;&#123;&#100;&#105;&#109;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#82;&#125;&#41;&#44;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>where <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ab41b2e51b8afd5fbfb9ae4f9fd1e881_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#82;&#125;" title="Rendered by QuickLaTeX.com" height="12" width="15" style="vertical-align: 0px;"/> is some optional space of augmentations (e.g. translations/rotations). Thus, we can typically expect <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-4a48166dd2d2d629c4f8d493fbb01868_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#114;&#109;&#123;&#100;&#105;&#109;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#77;&#125;&#95;&#115;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="71" style="vertical-align: -5px;"/> to be much lower than <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-d865a64aa254d26f90c3f0c3bd85ac3b_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#114;&#109;&#123;&#100;&#105;&#109;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#83;&#125;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="55" style="vertical-align: -5px;"/>, and thus: <strong>it is very easy to find adversarial examples that hack any state to any other desired state.</strong></p>
<p>As a result, the dynamics optimization</p>
<p class="ql-center-displayed-equation" style="line-height: 52px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-b32171577f474230feb8469c32d3c3e4_l3.png" height="52" width="181" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#115;&#117;&#109;&#95;&#123;&#116;&#61;&#48;&#125;&#94;&#123;&#84;&#45;&#49;&#125;&#32;&#92;&#98;&#105;&#103;&#92;&#124;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#116;&#44;&#97;&#95;&#116;&#41;&#32;&#45;&#32;&#115;&#95;&#123;&#116;&#43;&#49;&#125;&#92;&#98;&#105;&#103;&#92;&#124;&#95;&#50;&#94;&#50;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>feels incredibly “sticky,” as the base points <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-e86795deb37ff5f5055e741b17eb25d7_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#115;&#95;&#116;" title="Rendered by QuickLaTeX.com" height="11" width="13" style="vertical-align: -3px;"/> can easily trick <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-c502880341922d69202070782fbc9ab3_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#70;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;" title="Rendered by QuickLaTeX.com" height="15" width="18" style="vertical-align: -3px;"/> into thinking it’s already made its local goal.<sup><a href="http://bair.berkeley.edu/blog/2026/04/20/grasp/#fn1" id="ref1" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">1</a></sup></p>
<p><figure style="text-align: center;">
  <img decoding="async" src="https://bair.berkeley.edu/static/blog/grasp/pusht_adversarial.gif" alt="Adversarial world model example" style="max-width: 70%;" /><br />
</figure>
</p>
<hr />
<div id="fn1" style="font-size: 0.88em; margin: 0.75em 0; padding-left: 1em; border-left: 3px solid #ddd; color: #5f5f5f;">
<p><strong>1.</strong> This adversarial robustness issue, while particularly bad for lifted-state approaches, is not unique to them. Even for serial optimization methods that optimize through the full rollout map <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-6b9827b4af48fd65e7e83838bfd859f6_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#70;&#125;&#94;&#84;" title="Rendered by QuickLaTeX.com" height="16" width="25" style="vertical-align: -1px;"/>, it is possible to get into unseen states, where it is very easy to have a normal component fed into the sensitive normal components of <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-0ff50bc1113354825aa0a8ee24984eeb_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#68;&#95;&#115;&#32;&#70;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;" title="Rendered by QuickLaTeX.com" height="15" width="40" style="vertical-align: -3px;"/>. The action Jacobian’s chain rule expansion is</p>
<p class="ql-center-displayed-equation" style="line-height: 52px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3ce71c40a5f7f6ce0dc00795502ef2c8_l3.png" height="52" width="243" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#66;&#105;&#103;&#108;&#40;&#92;&#112;&#114;&#111;&#100;&#95;&#123;&#116;&#61;&#49;&#125;&#94;&#84;&#32;&#68;&#95;&#115;&#32;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#116;&#44;&#32;&#97;&#95;&#116;&#41;&#92;&#66;&#105;&#103;&#114;&#41;&#32;&#68;&#95;&#123;&#97;&#95;&#48;&#125;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#48;&#44;&#32;&#97;&#95;&#48;&#41;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>See what happens if any stage of the product has any component normal to the data manifold. <a href="http://bair.berkeley.edu/blog/2026/04/20/grasp/#ref1" style="color: #4d6b92;" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" /></a></p>
</div>
<hr />
<h3 id="our-fix">Our fix</h3>
<p>This is where our new planner GRASP comes in. The main observation: while <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-0ff50bc1113354825aa0a8ee24984eeb_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#68;&#95;&#115;&#32;&#70;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;" title="Rendered by QuickLaTeX.com" height="15" width="40" style="vertical-align: -3px;"/> is untrustworthy and adversarial, the action space is usually low-dimensional and exhaustively trained, so <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-c2f1132e9862e1f0b7226021ef92e3e4_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#68;&#95;&#97;&#32;&#70;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;" title="Rendered by QuickLaTeX.com" height="15" width="41" style="vertical-align: -3px;"/> is actually reasonable to optimize through and doesn’t suffer from the adversarial robustness issue!</p>
<p><figure style="text-align: center;">
  <img decoding="async" src="https://bair.berkeley.edu/static/blog/grasp/network_diagram.jpg" alt="Network diagram showing high-dim state vs low-dim action" style="max-width: 65%;" /><figcaption><em>The action input is usually lower-dimensional and densely trained (the model has seen every action direction), so action gradients are much better behaved.</em></figcaption></figure>
</p>
<p>At its core, <strong>GRASP builds a first-order lifted state / collocation-based planner that is only dependent on action Jacobians through the world model.</strong> We thus exploit the differentiability of learned world models <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-c502880341922d69202070782fbc9ab3_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#70;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;" title="Rendered by QuickLaTeX.com" height="15" width="18" style="vertical-align: -3px;"/>, while not falling victim to the inherent sensitivity of the state Jacobians <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-0ff50bc1113354825aa0a8ee24984eeb_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#68;&#95;&#115;&#32;&#70;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;" title="Rendered by QuickLaTeX.com" height="15" width="40" style="vertical-align: -3px;"/>.</p>
<h2 id="grasp-gradient-relaxed-stochastic-planner">GRASP: Gradient <strong>RelAxed</strong> <strong>S</strong>tochastic <strong>P</strong>lanner</h2>
<p>As noted before, we start with the collocation planning objective, where we lift the states and relax dynamics into a penalty:</p>
<p class="ql-center-displayed-equation" style="line-height: 52px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-863f11a3d6371cdc4342477c54c6f78f_l3.png" height="52" width="510" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#109;&#105;&#110;&#95;&#123;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#115;&#125;&#44;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#125;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#115;&#125;&#44;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41;&#32;&#61;&#32;&#92;&#115;&#117;&#109;&#95;&#123;&#116;&#61;&#48;&#125;&#94;&#123;&#84;&#45;&#49;&#125;&#32;&#92;&#98;&#105;&#103;&#92;&#124;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#95;&#116;&#44;&#97;&#95;&#116;&#41;&#32;&#45;&#32;&#115;&#95;&#123;&#116;&#43;&#49;&#125;&#92;&#98;&#105;&#103;&#92;&#124;&#95;&#50;&#94;&#50;&#44; &#92;&#113;&#117;&#97;&#100;&#32;&#92;&#116;&#101;&#120;&#116;&#123;&#119;&#105;&#116;&#104;&#32;&#125;&#32;&#115;&#95;&#48;&#32;&#92;&#116;&#101;&#120;&#116;&#123;&#32;&#102;&#105;&#120;&#101;&#100;&#32;&#97;&#110;&#100;&#32;&#125;&#32;&#115;&#95;&#84;&#61;&#103;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>We then make two key additions.</p>
<h2 id="ingredient-1-exploration-by-noising-the-state-iterates">Ingredient 1: Exploration by noising the <strong>state iterates</strong></h2>
<p>Even with a smoother objective, planning is nonconvex. We introduce exploration by injecting Gaussian noise into the <strong>virtual state updates</strong> during optimization.</p>
<p>A simple version:</p>
<p class="ql-center-displayed-equation" style="line-height: 19px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-75ff21e0a372181401e643f6a15de711_l3.png" height="19" width="339" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#115;&#95;&#116;&#32;&#92;&#108;&#101;&#102;&#116;&#97;&#114;&#114;&#111;&#119;&#32;&#115;&#95;&#116;&#32;&#45;&#32;&#92;&#101;&#116;&#97;&#95;&#115;&#32;&#92;&#110;&#97;&#98;&#108;&#97;&#95;&#123;&#115;&#95;&#116;&#125;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#32;&#43;&#32;&#92;&#115;&#105;&#103;&#109;&#97;&#95;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#115;&#116;&#97;&#116;&#101;&#125;&#125;&#32;&#92;&#120;&#105;&#44;&#32;&#92;&#113;&#113;&#117;&#97;&#100;&#32;&#92;&#120;&#105;&#92;&#115;&#105;&#109;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#78;&#125;&#40;&#48;&#44;&#73;&#41;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>Actions are still updated by non-stochastic descent:</p>
<p class="ql-center-displayed-equation" style="line-height: 16px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-cba1022bc8ad5d65da23bfac98316acb_l3.png" height="16" width="141" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#97;&#95;&#116;&#32;&#92;&#108;&#101;&#102;&#116;&#97;&#114;&#114;&#111;&#119;&#32;&#97;&#95;&#116;&#32;&#45;&#32;&#92;&#101;&#116;&#97;&#95;&#97;&#32;&#92;&#110;&#97;&#98;&#108;&#97;&#95;&#123;&#97;&#95;&#116;&#125;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>The state noise helps you “hop” between basins in the lifted space, while the actions remain guided by gradients. We found that specifically noising states here (as opposed to actions) finds a good balance of exploration and the ability to find sharper minima.<sup><a href="http://bair.berkeley.edu/blog/2026/04/20/grasp/#fn2" id="ref2" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">2</a></sup></p>
<hr />
<div id="fn2" style="font-size: 0.88em; margin: 0.75em 0; padding-left: 1em; border-left: 3px solid #ddd; color: #5f5f5f;">
<p><strong>2.</strong> Because we only noise the states (and not the actions), the corresponding dynamics are not truly Langevin dynamics. <a href="http://bair.berkeley.edu/blog/2026/04/20/grasp/#ref2" style="color: #4d6b92;" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/21a9.png" alt="↩" class="wp-smiley" style="height: 1em; max-height: 1em;" /></a></p>
</div>
<hr />
<h2 id="ingredient-2-reshape-gradients-stop-brittle-state-input-gradients-keep-action-gradients">Ingredient 2: Reshape gradients: stop brittle state-input gradients, keep action gradients</h2>
<p>As discussed, the fragile pathway is the gradient that flows <em>into the state input</em> of the world model, <span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-0ff50bc1113354825aa0a8ee24984eeb_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#68;&#95;&#115;&#32;&#70;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;" title="Rendered by QuickLaTeX.com" height="15" width="40" style="vertical-align: -3px;"/></span>. The most straightforward way to do this initially is to just stop state gradients into <span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-c502880341922d69202070782fbc9ab3_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#70;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;" title="Rendered by QuickLaTeX.com" height="15" width="18" style="vertical-align: -3px;"/></span> directly:</p>
<ul>
<li>Let <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-d6a08548d66ef341e02d54f6937ad037_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#98;&#97;&#114;&#123;&#115;&#125;&#95;&#116;" title="Rendered by QuickLaTeX.com" height="14" width="13" style="vertical-align: -3px;"/> be the same value as <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-e86795deb37ff5f5055e741b17eb25d7_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#115;&#95;&#116;" title="Rendered by QuickLaTeX.com" height="11" width="13" style="vertical-align: -3px;"/>, but with gradients stopped.</li>
</ul>
<p>Define the <strong>stop-gradient dynamics loss</strong>:</p>
<p class="ql-center-displayed-equation" style="line-height: 52px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-cd8e96c9d63f8a87db71c295647be60c_l3.png" height="52" width="284" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#95;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#100;&#121;&#110;&#125;&#125;&#94;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#115;&#103;&#125;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#115;&#125;&#44;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41; &#61;&#32;&#92;&#115;&#117;&#109;&#95;&#123;&#116;&#61;&#48;&#125;&#94;&#123;&#84;&#45;&#49;&#125;&#32;&#92;&#98;&#105;&#103;&#92;&#124;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#92;&#98;&#97;&#114;&#123;&#115;&#125;&#95;&#116;&#44;&#32;&#97;&#95;&#116;&#41;&#32;&#45;&#32;&#115;&#95;&#123;&#116;&#43;&#49;&#125;&#92;&#98;&#105;&#103;&#92;&#124;&#95;&#50;&#94;&#50;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>This alone does not work. Notice now states only follow the previous state’s step, without anything forcing the base states to chase the next ones. As a result, there are trivial minima for just stopping at the origin, then only for the final action trying to get to the goal in one step.</p>
<h3 id="dense-goal-shaping">Dense goal shaping</h3>
<p>We can view the above issue as the goal’s signal being cut off entirely from previous states. One way to fix this is to simply add a dense goal term throughout prediction:</p>
<p class="ql-center-displayed-equation" style="line-height: 52px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-8bf9d85259fdc485efc9660032128d29_l3.png" height="52" width="263" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#95;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#103;&#111;&#97;&#108;&#125;&#125;&#94;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#115;&#103;&#125;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#115;&#125;&#44;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41; &#61;&#32;&#92;&#115;&#117;&#109;&#95;&#123;&#116;&#61;&#48;&#125;&#94;&#123;&#84;&#45;&#49;&#125;&#32;&#92;&#98;&#105;&#103;&#92;&#124;&#70;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#92;&#98;&#97;&#114;&#123;&#115;&#125;&#95;&#116;&#44;&#32;&#97;&#95;&#116;&#41;&#32;&#45;&#32;&#103;&#92;&#98;&#105;&#103;&#92;&#124;&#95;&#50;&#94;&#50;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>In normal settings this would over-bias towards the greedy solution of straight chasing the goal, but this is balanced in our setting by the stop-gradient dynamics loss’s bias towards feasible dynamics. The final objective is then as follows:</p>
<p class="ql-center-displayed-equation" style="line-height: 24px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-d1f4f133db4d9c011c84ac420e31def9_l3.png" height="24" width="267" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#115;&#125;&#44;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41;&#32;&#61;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#95;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#100;&#121;&#110;&#125;&#125;&#94;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#115;&#103;&#125;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#115;&#125;&#44;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41;&#32;&#43;&#32;&#92;&#103;&#97;&#109;&#109;&#97;&#32;&#92;&#44;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#95;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#103;&#111;&#97;&#108;&#125;&#125;&#94;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#115;&#103;&#125;&#125;&#40;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#115;&#125;&#44;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>The result is a planning optimization objective that does not have dependence on state gradients.</p>
<hr />
<h2 id="periodic-sync-briefly-return-to-true-rollout-gradients">Periodic “sync”: briefly return to true rollout gradients</h2>
<p>The lifted stop-gradient objective is great for <strong>fast, guided exploration</strong>, but it’s still an approximation of the original serial rollout objective.</p>
<p>So every <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-0efc4f0539590861f708da095a045355_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#75;&#95;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#115;&#121;&#110;&#99;&#125;&#125;" title="Rendered by QuickLaTeX.com" height="18" width="41" style="vertical-align: -6px;"/> iterations, GRASP does a short refinement phase:</p>
<ol>
<li>Roll out from <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-7730c5a8c9bd4b8291e5231082f6b9c6_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#115;&#95;&#48;" title="Rendered by QuickLaTeX.com" height="11" width="15" style="vertical-align: -3px;"/> using current actions <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-91d6e96c2498d9d1364efcb44c9f1efe_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;" title="Rendered by QuickLaTeX.com" height="9" width="10" style="vertical-align: -1px;"/>, and take a few small gradient steps on the original serial loss:</li>
</ol>
<p class="ql-center-displayed-equation" style="line-height: 22px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-c542432dde5b7f9184dd1c220097eea2_l3.png" height="22" width="237" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#32;&#92;&#108;&#101;&#102;&#116;&#97;&#114;&#114;&#111;&#119;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#32;&#45;&#32;&#92;&#101;&#116;&#97;&#95;&#123;&#92;&#116;&#101;&#120;&#116;&#123;&#115;&#121;&#110;&#99;&#125;&#125;&#92;&#44;&#92;&#110;&#97;&#98;&#108;&#97;&#95;&#123;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#125;&#92;&#44;&#92;&#124;&#115;&#95;&#84;&#40;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#41;&#45;&#103;&#92;&#124;&#95;&#50;&#94;&#50;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>The lifted-state optimization still provides the core of the optimization, while this refinement step adds some assistance to keep states and actions grounded towards real trajectories. This refinement step can of course be replaced with a serial planner of your choice (e.g. CEM); the core idea is to still get some of the benefit of the full-path synchronization of serial planners, while still mostly using the benefits of the lifted-state planning.</p>
<hr />
<h2 id="how-grasp-addresses-long-range-planning">How GRASP addresses long-range planning</h2>
<p>Collocation-based planners offer a natural fix for long-horizon planning, but this optimization is quite difficult through modern world models due to adversarial robustness issues. <em>GRASP proposes a simple solution for a smoother collocation-based planner, alongside stable stochasticity for exploration</em>. As a result, longer-horizon planning ends up not only succeeding more, but also finding such successes faster:</p>
<p><figure style="text-align: center; margin: 1.25em 0;">
  <img decoding="async" src="https://bair.berkeley.edu/static/blog/grasp/pusht_zoomout.gif" alt="Push-T planning demo" style="max-width: 90%; height: auto;" /><figcaption style="font-size: 0.95em; margin-top: 0.5em;"><em>Push-T demo: longer-horizon planning with GRASP.</em></figcaption></figure>
</p>
<div class="grasp-results-table" style="overflow-x: auto; margin: 1em 0;">
<table>
<thead>
<tr>
<th>Horizon</th>
<th>CEM</th>
<th>GD</th>
<th>LatCo</th>
<th><strong>GRASP</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td>H=40</td>
<td><strong>61.4%</strong> / 35.3s</td>
<td>51.0% / 18.0s</td>
<td>15.0% / 598.0s</td>
<td>59.0% / <strong>8.5s</strong></td>
</tr>
<tr>
<td>H=50</td>
<td>30.2% / 96.2s</td>
<td>37.6% / 76.3s</td>
<td>4.2% / 1114.7s</td>
<td><strong>43.4%</strong> / <strong>15.2s</strong></td>
</tr>
<tr>
<td>H=60</td>
<td>7.2% / 83.1s</td>
<td>16.4% / 146.5s</td>
<td>2.0% / 231.5s</td>
<td><strong>26.2%</strong> / <strong>49.1s</strong></td>
</tr>
<tr>
<td>H=70</td>
<td>7.8% / 156.1s</td>
<td>12.0% / 103.1s</td>
<td>0.0% / —</td>
<td><strong>16.0%</strong> / <strong>79.9s</strong></td>
</tr>
<tr>
<td>H=80</td>
<td>2.8% / 132.2s</td>
<td>6.4% / 161.3s</td>
<td>0.0% / —</td>
<td><strong>10.4%</strong> / <strong>58.9s</strong></td>
</tr>
</tbody>
</table>
</div>
<p style="text-align: center; margin-top: 0.75em;"><em>Push-T results. Success rate (%) / median time to success. Bold = best in row. Note the median success time will bias higher with higher success rate; GRASP manages to be faster despite higher success rate.</em></p>
<hr />
<h2 id="whats-next">What’s next?</h2>
<p>There is still plenty of work to be done for modern world model planners. We want to exploit the gradient structure of learned world models, and collocation (lifted-state optimization) is a natural approach for long-horizon planning, but it’s crucial to understand typical gradient structure here: smooth and informative action gradients and brittle state gradients. We view GRASP as an initial iteration for such planners.</p>
<p>Extension to diffusion-based world models (deeper latent timesteps can be viewed as smoothed versions of the world model itself), more sophisticated optimizers and noising strategies, and integrating GRASP into either a closed-loop system or RL policy learning for adaptive long-horizon planning are all natural and interesting next steps.</p>
<p>I do genuinely think it’s an exciting time to be working on world model planners. It’s a funny sweet spot where the background literature (planning and control overall) is incredibly mature and well-developed, but the current setting (pure planning optimization over modern, large-scale world models) is still heavily underexplored. But, once we figure out all the right ideas, world model planners will likely become as commonplace as RL.</p>
<hr />
<p>For more details, read the <a href="https://arxiv.org/pdf/2602.00475" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">full paper</a> or visit the <a href="https://www.michaelpsenka.io/grasp/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">project website</a>.</p>
<hr />
<h2 id="citation">Citation</h2>
<div class="language-bibtex highlighter-rouge">
<div class="highlight">
<pre class="highlight"><code><span class="nc">@article</span><span class="p">{</span><span class="nl">psenka2026grasp</span><span class="p">,</span>
  <span class="na">title</span><span class="p">=</span><span class="s">{Parallel Stochastic Gradient-Based Planning for World Models}</span><span class="p">,</span>
  <span class="na">author</span><span class="p">=</span><span class="s">{Michael Psenka and Michael Rabbat and Aditi Krishnapriyan and Yann LeCun and Amir Bar}</span><span class="p">,</span>
  <span class="na">year</span><span class="p">=</span><span class="s">{2026}</span><span class="p">,</span>
  <span class="na">eprint</span><span class="p">=</span><span class="s">{2602.00475}</span><span class="p">,</span>
  <span class="na">archivePrefix</span><span class="p">=</span><span class="s">{arXiv}</span><span class="p">,</span>
  <span class="na">primaryClass</span><span class="p">=</span><span class="s">{cs.LG}</span><span class="p">,</span>
  <span class="na">url</span><span class="p">=</span><span class="s">{https://arxiv.org/abs/2602.00475}</span>
<span class="p">}</span>
</code></pre>
</div>
</div>
<hr />
<p>This article was initially published on the <a href="https://bair.berkeley.edu/blog/2026/04/20/grasp/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Goal representations for instruction following</title>
		<link>https://robohub.org/goal-representations-for-instruction-following/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sun, 22 Oct 2023 08:35:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://bair.berkeley.edu/blog/2023/10/17/grif/</guid>

					<description><![CDATA[











Goal Representations for Instruction Following






A longstanding goal of the field of robot learning has been to create generalist agents that can perform tasks for humans. Natural language has the potential to be an easy-to-use interfa...]]></description>
										<content:encoded><![CDATA[<img decoding="async" src="https://robohub.org/wp-content/uploads/2023/10/thumbnail.png" alt="" width="2842" height="1870" class="alignnone size-full wp-image-208539" srcset="https://robohub.org/wp-content/uploads/2023/10/thumbnail.png 2842w, https://robohub.org/wp-content/uploads/2023/10/thumbnail-425x280.png 425w, https://robohub.org/wp-content/uploads/2023/10/thumbnail-1024x674.png 1024w, https://robohub.org/wp-content/uploads/2023/10/thumbnail-768x505.png 768w, https://robohub.org/wp-content/uploads/2023/10/thumbnail-1536x1011.png 1536w, https://robohub.org/wp-content/uploads/2023/10/thumbnail-2048x1348.png 2048w" sizes="(max-width: 2842px) 100vw, 2842px" />
<p><strong>By Andre He, Vivek Myers</strong></p>
<p>A longstanding goal of the field of robot learning has been to create generalist agents that can perform tasks for humans. Natural language has the potential to be an easy-to-use interface for humans to specify arbitrary tasks, but it is difficult to train robots to follow language instructions. Approaches like language-conditioned behavioral cloning (LCBC) train policies to directly imitate expert actions conditioned on language, but require humans to annotate all training trajectories and generalize poorly across scenes and behaviors. Meanwhile, recent goal-conditioned approaches perform much better at general manipulation tasks, but do not enable easy task specification for human operators. How can we reconcile the ease of specifying tasks through LCBC-like approaches with the performance improvements of goal-conditioned learning?</p>
<p>Conceptually, an instruction-following robot requires two capabilities. It needs to ground the language instruction in the physical environment, and then be able to carry out a sequence of actions to complete the intended task. These capabilities do not need to be learned end-to-end from human-annotated trajectories alone, but can instead be learned separately from the appropriate data sources. Vision-language data from non-robot sources can help learn language grounding with generalization to diverse instructions and visual scenes. Meanwhile, unlabeled robot trajectories can be used to train a robot to reach specific goal states, even when they are not associated with language instructions.</p>
<p>Conditioning on visual goals (i.e. goal images) provides complementary benefits for policy learning. As a form of task specification, goals are desirable for scaling because they can be freely generated hindsight relabeling (any state reached along a trajectory can be a goal). This allows policies to be trained via goal-conditioned behavioral cloning (GCBC) on large amounts of unannotated and unstructured trajectory data, including data collected autonomously by the robot itself. Goals are also easier to ground since, as images, they can be directly compared pixel-by-pixel with other states.</p>
<p>However, goals are less intuitive for human users than natural language. In most cases, it is easier for a user to describe the task they want performed than it is to provide a goal image, which would likely require performing the task anyways to generate the image. By exposing a language interface for goal-conditioned policies, we can combine the strengths of both goal- and language- task specification to enable generalist robots that can be easily commanded. Our method, discussed below, exposes such an interface to generalize to diverse instructions and scenes using vision-language data, and improve its physical skills by digesting large unstructured robot datasets.</p>
<h2 id="goal-representations-for-instruction-following">Goal representations for instruction following</h2>
<div id="attachment_208540" style="width: 3750px" class="wp-caption alignnone"><img decoding="async" aria-describedby="caption-attachment-208540" src="https://robohub.org/wp-content/uploads/2023/10/figure1.png" alt="" width="3740" height="2094" class="size-full wp-image-208540" srcset="https://robohub.org/wp-content/uploads/2023/10/figure1.png 3740w, https://robohub.org/wp-content/uploads/2023/10/figure1-425x238.png 425w, https://robohub.org/wp-content/uploads/2023/10/figure1-1024x573.png 1024w, https://robohub.org/wp-content/uploads/2023/10/figure1-768x430.png 768w, https://robohub.org/wp-content/uploads/2023/10/figure1-1536x860.png 1536w, https://robohub.org/wp-content/uploads/2023/10/figure1-2048x1147.png 2048w" sizes="(max-width: 3740px) 100vw, 3740px" /><p id="caption-attachment-208540" class="wp-caption-text">The GRIF model consists of a language encoder, a goal encoder, and a policy network. The encoders respectively map language instructions and goal images into a shared task representation space, which conditions the policy network when predicting actions. The model can effectively be conditioned on either language instructions or goal images to predict actions, but we are primarily using goal-conditioned training as a way to improve the language-conditioned use case.</p></div>
<p>Our approach, <b>Goal Representations for Instruction Following (GRIF)</b>, jointly trains a language- and a goal- conditioned policy with aligned task representations. Our key insight is that these representations, aligned across language and goal modalities, enable us to effectively combine the benefits of goal-conditioned learning with a language-conditioned policy. The learned policies are then able to generalize across language and scenes after training on mostly unlabeled demonstration data.</p>
<p>We trained GRIF on a version of the <a href="https://rail-berkeley.github.io/bridgedata/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Bridge-v2 dataset</a> containing 7k labeled demonstration trajectories and 47k unlabeled ones within a kitchen manipulation setting. Since all the trajectories in this dataset had to be manually annotated by humans, being able to directly use the 47k trajectories without annotation significantly improves efficiency.</p>
<p>To learn from both types of data, GRIF is trained jointly with language-conditioned behavioral cloning (LCBC) and goal-conditioned behavioral cloning (GCBC). The labeled dataset contains both language and goal task specifications, so we use it to supervise both the language- and goal-conditioned predictions (i.e. LCBC and GCBC). The unlabeled dataset contains only goals and is used for GCBC. The difference between LCBC and GCBC is just a matter of selecting the task representation from the corresponding encoder, which is passed into a shared policy network to predict actions.</p>
<p>By sharing the policy network, we can expect some improvement from using the unlabeled dataset for goal-conditioned training. However,GRIF enables much stronger transfer between the two modalities by recognizing that some language instructions and goal images specify the same behavior. In particular, we exploit this structure by requiring that language- and goal- representations be similar for the same semantic task. Assuming this structure holds, unlabeled data can also benefit the language-conditioned policy since the goal representation approximates that of the missing instruction.</p>
<h2 id="alignment-through-contrastive-learning">Alignment through contrastive learning</h2>
<div id="attachment_208541" style="width: 1516px" class="wp-caption alignnone"><img decoding="async" aria-describedby="caption-attachment-208541" src="https://robohub.org/wp-content/uploads/2023/10/contrast.png" alt="" width="1506" height="1338" class="size-full wp-image-208541" srcset="https://robohub.org/wp-content/uploads/2023/10/contrast.png 1506w, https://robohub.org/wp-content/uploads/2023/10/contrast-425x378.png 425w, https://robohub.org/wp-content/uploads/2023/10/contrast-1024x910.png 1024w, https://robohub.org/wp-content/uploads/2023/10/contrast-768x682.png 768w" sizes="(max-width: 1506px) 100vw, 1506px" /><p id="caption-attachment-208541" class="wp-caption-text">We explicitly align representations between goal-conditioned and language-conditioned tasks on the labeled dataset through contrastive learning.</p></div>
<p>Since language often describes relative change, we choose to align representations of state-goal pairs with the language instruction (as opposed to just goal with language). Empirically, this also makes the representations easier to learn since they can omit most information in the images and focus on the change from state to goal.</p>
<p>We learn this alignment structure through an infoNCE objective on instructions and images from the labeled dataset. We train dual image and text encoders by doing contrastive learning on matching pairs of language and goal representations. The objective encourages high similarity between representations of the same task and low similarity for others, where the negative examples are sampled from other trajectories.</p>
<p>When using naive negative sampling (uniform from the rest of the dataset), the learned representations often ignored the actual task and simply aligned instructions and goals that referred to the same scenes. To use the policy in the real world, it is not very useful to associate language with a scene; rather we need it to disambiguate between different tasks in the same scene. Thus, we use a hard negative sampling strategy, where up to half the negatives are sampled from different trajectories in the same scene.</p>
<p>Naturally, this contrastive learning setup teases at pre-trained vision-language models like CLIP. They demonstrate effective zero-shot and few-shot generalization capability for vision-language tasks, and offer a way to incorporate knowledge from internet-scale pre-training. However, most vision-language models are designed for aligning a single static image with its caption without the ability to understand changes in the environment, and they perform poorly when having to pay attention to a single object in cluttered scenes.</p>
<p>To address these issues, we devise a mechanism to accommodate and fine-tune CLIP for aligning task representations. We modify the CLIP architecture so that it can operate on a pair of images combined with early fusion (stacked channel-wise). This turns out to be a capable initialization for encoding pairs of state and goal images, and one which is particularly good at preserving the pre-training benefits from CLIP.</p>
<h2 id="robot-policy-results">Robot policy results</h2>
<p>For our main result, we evaluate the GRIF policy in the real world on 15 tasks across 3 scenes. The instructions are chosen to be a mix of ones that are well-represented in the training data and novel ones that require some degree of compositional generalization. One of the scenes also features an unseen combination of objects.</p>
<p>We compare GRIF against plain LCBC and stronger baselines inspired by prior work like <a href="https://language-play.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">LangLfP</a> and <a href="https://sites.google.com/view/bc-z/home" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BC-Z</a>. LLfP corresponds to jointly training with LCBC and GCBC. BC-Z is an adaptation of the namesake method to our setting, where we train on LCBC, GCBC, and a simple alignment term. It optimizes the cosine distance loss between the task representations and does not use image-language pre-training.</p>
<p>The policies were susceptible to two main failure modes. They can fail to understand the language instruction, which results in them attempting another task or performing no useful actions at all. When language grounding is not robust, policies might even start an unintended task after having done the right task, since the original instruction is out of context.</p>
<p style="text-align: center;"><i>Examples of grounding failures</i></p>
<div style="display: flex; justify-content: space-between;">
<div style="width: 25%; text-align: center;">
        <img decoding="async" src="https://bair.berkeley.edu/static/blog/grif/grounding1.gif" alt="grounding failure 1" style="max-width: 100%; height: auto;" /></p>
<p width="80%" style="text-align:center; margin-left:10%; margin-right:10%; padding-bottom: -10px"><i>&#8220;put the mushroom in the metal pot&#8221;</i></p>
</p></div>
<div style="width: 25%; text-align: center;">
        <img decoding="async" src="https://bair.berkeley.edu/static/blog/grif/grounding2.gif" alt="grounding failure 2" style="max-width: 100%; height: auto;" /></p>
<p width="80%" style="text-align:center; margin-left:10%; margin-right:10%; padding-bottom: -10px"><i>&#8220;put the spoon on the towel&#8221;</i></p>
</p></div>
<div style="width: 25%; text-align: center;">
        <img decoding="async" src="https://bair.berkeley.edu/static/blog/grif/grounding3.gif" alt="grounding failure 3" style="max-width: 100%; height: auto;" /></p>
<p width="80%" style="text-align:center; margin-left:10%; margin-right:10%; padding-bottom: -10px"><i>&#8220;put the yellow bell pepper on the cloth&#8221;</i></p>
</p></div>
<div style="width: 25%; text-align: center;">
        <img decoding="async" src="https://bair.berkeley.edu/static/blog/grif/grounding4.gif" alt="grounding failure 4" style="max-width: 100%; height: auto;" /></p>
<p width="80%" style="text-align:center; margin-left:10%; margin-right:10%; padding-bottom: -10px"><i>&#8220;put the yellow bell pepper on the cloth&#8221;</i></p>
</p></div>
</div>
<p>The other failure mode is failing to manipulate objects. This can be due to missing a grasp, moving imprecisely, or releasing objects at the incorrect time. We note that these are not inherent shortcomings of the robot setup, as a GCBC policy trained on the entire dataset can consistently succeed in manipulation. Rather, this failure mode generally indicates an ineffectiveness in leveraging goal-conditioned data.</p>
<p style="text-align: center;"><i>Examples of manipulation failures</i></p>
<div style="display: flex; justify-content: space-between;">
<div style="width: 25%; text-align: center;">
        <img decoding="async" src="https://bair.berkeley.edu/static/blog/grif/manipulation1.gif" alt="manipulation failure 1" style="max-width: 100%; height: auto;" /></p>
<p width="80%" style="text-align:center; margin-left:10%; margin-right:10%; padding-bottom: -10px"><i>&#8220;move the bell pepper to the left of the table&#8221;</i></p>
</p></div>
<div style="width: 25%; text-align: center;">
        <img decoding="async" src="https://bair.berkeley.edu/static/blog/grif/manipulation2.gif" alt="manipulation failure 2" style="max-width: 100%; height: auto;" /></p>
<p width="80%" style="text-align:center; margin-left:10%; margin-right:10%; padding-bottom: -10px"><i>&#8220;put the bell pepper in the pan&#8221;</i></p>
</p></div>
<div style="width: 25%; text-align: center;">
        <img decoding="async" src="https://bair.berkeley.edu/static/blog/grif/manipulation3.gif" alt="manipulation failure 3" style="max-width: 100%; height: auto;" /></p>
<p width="80%" style="text-align:center; margin-left:10%; margin-right:10%; padding-bottom: -10px"><i>&#8220;move the towel next to the microwave&#8221;</i></p>
</p></div>
</div>
<p>Comparing the baselines, they each suffered from these two failure modes to different extents. LCBC relies solely on the small labeled trajectory dataset, and its poor manipulation capability prevents it from completing any tasks. LLfP jointly trains the policy on labeled and unlabeled data and shows significantly improved manipulation capability from LCBC. It achieves reasonable success rates for common instructions, but fails to ground more complex instructions. BC-Z’s alignment strategy also improves manipulation capability, likely because alignment improves the transfer between modalities. However, without external vision-language data sources, it still struggles to generalize to new instructions.</p>
<p>GRIF shows the best generalization while also having strong manipulation capabilities. It is able to ground the language instructions and carry out the task even when many distinct tasks are possible in the scene. We show some rollouts and the corresponding instructions below.</p>
<p style="text-align: center;"><i>Policy Rollouts from GRIF</i></p>
<div style="display: flex; justify-content: space-between;">
<div style="width: 25%; text-align: center;">
        <img decoding="async" src="https://bair.berkeley.edu/static/blog/grif/grif1.gif" alt="rollout 1" style="max-width: 100%; height: auto;" /></p>
<p width="80%" style="text-align:center; margin-left:10%; margin-right:10%; padding-bottom: -10px"><i>&#8220;move the pan to the front&#8221;</i></p>
</p></div>
<div style="width: 25%; text-align: center;">
        <img decoding="async" src="https://bair.berkeley.edu/static/blog/grif/grif2.gif" alt="rollout 2" style="max-width: 100%; height: auto;" /></p>
<p width="80%" style="text-align:center; margin-left:10%; margin-right:10%; padding-bottom: -10px"><i>&#8220;put the bell pepper in the pan&#8221;</i></p>
</p></div>
<div style="width: 25%; text-align: center;">
        <img decoding="async" src="https://bair.berkeley.edu/static/blog/grif/grif3.gif" alt="rollout 3" style="max-width: 100%; height: auto;" /></p>
<p width="80%" style="text-align:center; margin-left:10%; margin-right:10%; padding-bottom: -10px"><i>&#8220;put the knife on the purple cloth&#8221;</i></p>
</p></div>
<div style="width: 25%; text-align: center;">
        <img decoding="async" src="https://bair.berkeley.edu/static/blog/grif/grif4.gif" alt="rollout 4" style="max-width: 100%; height: auto;" /></p>
<p width="80%" style="text-align:center; margin-left:10%; margin-right:10%; padding-bottom: -10px"><i>&#8220;put the spoon on the towel&#8221;</i></p>
</p></div>
</div>
<h2 id="conclusion">Conclusion</h2>
<p>GRIF enables a robot to utilize large amounts of unlabeled trajectory data to learn goal-conditioned policies, while providing a “language interface” to these policies via aligned language-goal task representations. In contrast to prior language-image alignment methods, our representations align changes in state to language, which we show leads to significant improvements over standard CLIP-style image-language alignment objectives. Our experiments demonstrate that our approach can effectively leverage unlabeled robotic trajectories, with large improvements in performance over baselines and methods that only use the language-annotated data</p>
<p>Our method has a number of limitations that could be addressed in future work. GRIF is not well-suited for tasks where instructions say more about how to do the task than what to do (e.g., “pour the water slowly”)—such qualitative instructions might require other types of alignment losses that consider the intermediate steps of task execution. GRIF also assumes that all language grounding comes from the portion of our dataset that is fully annotated or a pre-trained VLM. An exciting direction for future work would be to extend our alignment loss to utilize human video data to learn rich semantics from Internet-scale data. Such an approach could then use this data to improve grounding on language outside the robot dataset and enable broadly generalizable robot policies that can follow user instructions.</p>
<hr />
<p>This post is based on the following paper:</p>
<ul>
<li>
    <a href="https://arxiv.org/abs/2307.00117" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"><strong>Goal Representations for Instruction Following: A Semi-Supervised Language Interface to Control</strong></a><br />
    <a href="https://people.eecs.berkeley.edu/~vmyers/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Vivek&nbsp;Myers</a>*, Andre&nbsp;He*, <a href="https://kuanfang.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Kuan&nbsp;Fang</a>, <a href="https://homerwalke.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Homer&nbsp;Walke</a>, Philippe&nbsp;Hansen-Estruch, <a href="https://www.chinganc.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Ching-An&nbsp;Cheng</a>, <a href="https://mihaij.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Mihai&nbsp;Jalobeanu</a>, <a href="https://www.microsoft.com/en-us/research/people/akolobov/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Andrey&nbsp;Kolobov</a>, <a href="http://people.eecs.berkeley.edu/~anca/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Anca&nbsp;Dragan</a>, and <a href="https://people.eecs.berkeley.edu/~svlevine/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sergey&nbsp;Levine</a>
    </li>
</ul>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Interactive fleet learning</title>
		<link>https://robohub.org/interactive-fleet-learning/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 12 Apr 2023 14:55:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://bair.berkeley.edu/blog/2023/04/06/ifl/</guid>

					<description><![CDATA[












Figure 1: “Interactive Fleet Learning” (IFL) refers to robot fleets in industry and academia that fall back on human teleoperators when necessary and continually learn from them over time.


In the last few years we have seen an excitin...]]></description>
										<content:encoded><![CDATA[<div id="attachment_20703688" style="width: 650px" class="wp-caption alignnone"><img decoding="async" aria-describedby="caption-attachment-20703688" src="https://bair.berkeley.edu/static/blog/ifl/figure2.jpg" alt="" width="640" height="452" class="size-full wp-image-20703688" /><p id="caption-attachment-20703688" class="wp-caption-text">Commercial and industrial deployments of robot fleets: package delivery (top left), food delivery (bottom left), e-commerce order fulfillment at Ambi Robotics (top right), autonomous taxis at Waymo (bottom right).</p></div>
<p>In the last few years we have seen an exciting development in robotics and artificial intelligence: large fleets of robots have left the lab and entered the real world. <a href="https://waymo.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Waymo</a>, for example, has over 700 self-driving cars operating in Phoenix and San Francisco and is <a href="https://blog.waymo.com/2022/10/next-stop-for-waymo-one-los-angeles.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">currently expanding to Los Angeles</a>. Other industrial deployments of robot fleets include applications like e-commerce order fulfillment at <a href="https://www.amazon.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Amazon</a> and <a href="https://www.ambirobotics.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Ambi Robotics</a> as well as food delivery at <a href="https://www.nuro.ai/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Nuro</a> and <a href="https://www.kiwibot.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Kiwibot</a>.</p>
<div id="attachment_20703699" style="width: 650px" class="wp-caption alignnone"><img decoding="async" aria-describedby="caption-attachment-20703699" src="https://bair.berkeley.edu/static/blog/ifl/figure1.gif" alt="" width="640" height="452" class="size-full wp-image-20703699" /><p id="caption-attachment-20703699" class="wp-caption-text">Figure 1: “Interactive Fleet Learning” (IFL) refers to robot fleets in industry and academia that fall back on human teleoperators when necessary and continually learn from them over time.</p></div>
<p>These robots use recent advances in deep learning to operate autonomously in unstructured environments. By pooling data from all robots in the fleet, the entire fleet can efficiently learn from the experience of each individual robot. Furthermore, due to advances in <a href="https://ieeexplore.ieee.org/document/7006734" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">cloud robotics</a>, the fleet can offload data, memory, and computation (e.g., training of large models) to the cloud via the Internet. This approach is known as “Fleet Learning,” a term popularized by Elon Musk in <a href="https://electrek.co/2016/09/11/transcript-elon-musks-press-conference-about-tesla-autopilot-under-v8-0-update-part-1/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">2016 press releases about Tesla Autopilot</a> and used in press communications by <a href="https://www.tri.global/news/tri-teaching-robots-help-people-their-homes" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Toyota Research Institute</a>, <a href="https://wayve.ai/technology/fleet-learning-technology/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wayve AI</a>, and others. A robot fleet is a modern analogue of a fleet of ships, where the word <em>fleet</em> has an etymology tracing back to <em>flēot</em> (‘ship’) and <em>flēotan</em> (‘float’) in Old English.</p>
<p>Data-driven approaches like fleet learning, however, face the problem of the <a href="https://www.forbes.com/sites/lanceeliot/2021/07/13/whether-those-endless-edge-or-corner-cases-are-the-long-tail-doom-for-ai-self-driving-cars/?sh=573981be5933" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">“long tail”</a>: the robots inevitably encounter new scenarios and edge cases that are not represented in the dataset. Naturally, we can’t expect the future to be the same as the past! How, then, can these robotics companies ensure sufficient reliability for their services?</p>
<p>One answer is to fall back on remote humans over the Internet, who can interactively take control and “tele-operate” the system when the robot policy is unreliable during task execution. Teleoperation has a rich history in robotics: <a href="https://goldberg.berkeley.edu/pubs/Nature-Robots-and-Return-to-Collaborative-Intelligence.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">the world’s first robots were teleoperated</a> during WWII to handle radioactive materials, and the <a href="https://en.wikipedia.org/wiki/Telegarden" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Telegarden</a> pioneered robot control over the Internet in 1994. With continual learning, the human teleoperation data from these interventions can iteratively improve the robot policy and reduce the robots’ reliance on their human supervisors over time. Rather than a discrete jump to full robot autonomy, this strategy offers a continuous alternative that approaches full autonomy over time while simultaneously enabling reliability in robot systems <em>today</em>.</p>
<p>The use of human teleoperation as a fallback mechanism is increasingly popular in modern robotics companies: Waymo calls it <a href="https://www.theatlantic.com/technology/archive/2018/08/waymos-robot-cars-and-the-humans-who-tend-to-them/568051/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">“fleet response,”</a> Zoox calls it <a href="https://twitter.com/zoox/status/1415737908112203776" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">“TeleGuidance,”</a> and Amazon calls it <a href="https://www.amazon.science/latest-news/robin-deals-with-a-world-where-things-are-changing-all-around-it" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">“continual learning.”</a> Last year, a software platform for remote driving called <a href="https://phantom.auto/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Phantom Auto</a> was recognized by Time Magazine as one of their <a href="https://time.com/collection/best-inventions-2022/6224834/phantom-auto-remote-operation-platform-for-logistics/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Top 10 Inventions of 2022</a>. And just last month, <a href="https://www.therobotreport.com/john-deere-acquires-sparkais-human-in-the-loop-tech/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">John Deere acquired SparkAI</a>, a startup that develops software for resolving edge cases with humans in the loop.</p>
<div id="attachment_20703677" style="width: 650px" class="wp-caption alignnone"><img decoding="async" aria-describedby="caption-attachment-20703677" src="https://bair.berkeley.edu/static/blog/ifl/figure3.jpg" alt="" width="640" height="452" class="size-full wp-image-20703677" /><p id="caption-attachment-20703677" class="wp-caption-text">A remote human teleoperator at Phantom Auto, a software platform for enabling remote driving over the Internet.</p></div>
<p>Despite this growing trend in industry, however, there has been comparatively little focus on this topic in academia. As a result, robotics companies have had to rely on ad hoc solutions for determining when their robots should cede control. The closest analogue in academia is <a href="https://arxiv.org/abs/2211.00600" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">interactive imitation learning (IIL)</a>, a paradigm in which a robot intermittently cedes control to a human supervisor and learns from these interventions over time. There have been a number of IIL algorithms in recent years for the single-robot, single-human setting including <a href="https://arxiv.org/abs/1011.0686" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DAgger</a> and variants such as <a href="https://arxiv.org/abs/1810.02890" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">HG-DAgger</a>, <a href="https://arxiv.org/abs/1605.06450" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SafeDAgger</a>, <a href="https://arxiv.org/abs/1807.08364" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">EnsembleDAgger</a>, and <a href="https://arxiv.org/abs/2109.08273" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ThriftyDAgger</a>; nevertheless, when and how to switch between robot and human control is still an open problem. This is even less understood when the notion is generalized to robot fleets, with multiple robots and multiple human supervisors.</p>
<h2 id="ifl-formalism-and-algorithms">IFL Formalism and Algorithms</h2>
<p>To this end, in a <a href="https://proceedings.mlr.press/v205/hoque23a.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recent paper at the Conference on Robot Learning</a> we introduced the paradigm of <em>Interactive Fleet Learning (IFL)</em>, the first formalism in the literature for interactive learning with multiple robots and multiple humans. As we’ve seen that this phenomenon already occurs in industry, we can now use the phrase “interactive fleet learning” as unified terminology for robot fleet learning that falls back on human control, rather than keep track of the names of every individual corporate solution (“fleet response”, “TeleGuidance”, etc.). IFL scales up robot learning with four key components:</p>
<ol>
<li><strong>On-demand supervision.</strong> Since humans cannot effectively monitor the execution of multiple robots at once and are prone to fatigue, the allocation of robots to humans in IFL is automated by some allocation policy <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-707fcec15e450815730425a6607a1858_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#111;&#109;&#101;&#103;&#97;" title="Rendered by QuickLaTeX.com" height="8" width="11" style="vertical-align: 0px;"/>. Supervision is requested “on-demand” by the robots rather than placing the burden of continuous monitoring on the humans.</li>
<li><strong>Fleet supervision.</strong> On-demand supervision enables effective allocation of limited human attention to large robot fleets. IFL allows the number of robots to significantly exceed the number of humans (e.g., by a factor of 10:1 or more).</li>
<li><strong>Continual learning.</strong> Each robot in the fleet can learn from its own mistakes as well as the mistakes of the other robots, allowing the amount of required human supervision to taper off over time.</li>
<li><strong>The Internet.</strong> Thanks to mature and ever-improving Internet technology, the human supervisors do not need to be physically present. Modern computer networks enable <a href="https://venturebeat.com/business/how-teleoperation-could-enable-remote-work-for-more-industries/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">real-time remote teleoperation</a> at vast distances.</li>
</ol>
<div id="attachment_20703666" style="width: 650px" class="wp-caption alignnone"><img decoding="async" aria-describedby="caption-attachment-20703666" src="https://bair.berkeley.edu/static/blog/ifl/figure4.jpg" alt="" width="640" height="452" class="size-full wp-image-20703666" /><p id="caption-attachment-20703666" class="wp-caption-text">In the Interactive Fleet Learning (IFL) paradigm, M humans are allocated to the robots that need the most help in a fleet of N robots (where N can be much larger than M). The robots share policy <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-1486f17cc0370e74cc821f0adc6acc88_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#112;&#105;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#95;&#116;&#125;" title="Rendered by QuickLaTeX.com" height="13" width="20" style="vertical-align: -5px;"/> and learn from human interventions over time.</p></div>
<p>We assume that the robots share a common control policy <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-1486f17cc0370e74cc821f0adc6acc88_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#112;&#105;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#95;&#116;&#125;" title="Rendered by QuickLaTeX.com" height="13" width="20" style="vertical-align: -5px;"/> and that the humans share a common control policy <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-b51b3997ff04cfb41ec071c0399be21d_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#112;&#105;&#95;&#72;" title="Rendered by QuickLaTeX.com" height="11" width="22" style="vertical-align: -3px;"/>. We also assume that the robots operate in independent environments with identical state and action spaces (but not identical states). Unlike a robot <em>swarm</em> of typically low-cost robots that coordinate to achieve a common objective in a shared environment, a robot <em>fleet</em> simultaneously executes a shared policy in distinct parallel environments (e.g., different bins on an assembly line).</p>
<p>The goal in IFL is to find an optimal supervisor allocation policy <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-707fcec15e450815730425a6607a1858_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#111;&#109;&#101;&#103;&#97;" title="Rendered by QuickLaTeX.com" height="8" width="11" style="vertical-align: 0px;"/>, a mapping from <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-42e999c68301fe978828dca7d5ac74eb_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#115;&#125;&#94;&#116;" title="Rendered by QuickLaTeX.com" height="15" width="13" style="vertical-align: 0px;"/> (the state of all robots at time <em>t</em>) and the shared policy <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-1486f17cc0370e74cc821f0adc6acc88_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#112;&#105;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#95;&#116;&#125;" title="Rendered by QuickLaTeX.com" height="13" width="20" style="vertical-align: -5px;"/> to a binary matrix that indicates which human will be assigned to which robot at time <em>t</em>. The IFL objective is a novel metric we call the “return on human effort” (ROHE):</p>
<p class="ql-center-displayed-equation" style="line-height: 54px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-860917029c4453032009999820787e2b_l3.png" height="54" width="363" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#091;&#92;&#109;&#97;&#120;&#95;&#123;&#92;&#111;&#109;&#101;&#103;&#97;&#32;&#92;&#105;&#110;&#32;&#92;&#79;&#109;&#101;&#103;&#97;&#125;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#69;&#125;&#95;&#123;&#92;&#116;&#97;&#117;&#32;&#92;&#115;&#105;&#109;&#32;&#112;&#95;&#123;&#92;&#111;&#109;&#101;&#103;&#97;&#44;&#32;&#92;&#116;&#104;&#101;&#116;&#97;&#95;&#48;&#125;&#40;&#92;&#116;&#97;&#117;&#41;&#125;&#32;&#92;&#108;&#101;&#102;&#116;&#091;&#92;&#102;&#114;&#97;&#99;&#123;&#77;&#125;&#123;&#78;&#125;&#32;&#92;&#99;&#100;&#111;&#116;&#32;&#92;&#102;&#114;&#97;&#99;&#123;&#92;&#115;&#117;&#109;&#95;&#123;&#116;&#61;&#48;&#125;&#94;&#84;&#32;&#92;&#98;&#97;&#114;&#123;&#114;&#125;&#40;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#115;&#125;&#94;&#116;&#44;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#94;&#116;&#41;&#125;&#123;&#49;&#43;&#92;&#115;&#117;&#109;&#95;&#123;&#116;&#61;&#48;&#125;&#94;&#84;&#32;&#92;&#124;&#92;&#111;&#109;&#101;&#103;&#97;&#40;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#115;&#125;&#94;&#116;&#44;&#32;&#92;&#112;&#105;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#95;&#116;&#125;&#44;&#32;&#92;&#99;&#100;&#111;&#116;&#41;&#32;&#92;&#124;&#94;&#50;&#32;&#95;&#70;&#125;&#32;&#92;&#114;&#105;&#103;&#104;&#116;&#093;&#92;&#093;" title="Rendered by QuickLaTeX.com"/></p>
<p>where the numerator is the total reward across robots and timesteps and the denominator is the total amount of human actions across robots and timesteps. Intuitively, the ROHE measures the performance of the fleet normalized by the total human supervision required. See the <a href="https://arxiv.org/abs/2206.14349" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a> for more of the mathematical details.</p>
<p>Using this formalism, we can now instantiate and compare IFL algorithms (i.e., allocation policies) in a principled way. We propose a family of IFL algorithms called Fleet-DAgger, where the policy learning algorithm is interactive imitation learning and each Fleet-DAgger algorithm is parameterized by a unique priority function <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-888d19958f4bde272de848ceaebc9630_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#104;&#97;&#116;&#32;&#112;&#58;&#32;&#40;&#115;&#44;&#32;&#92;&#112;&#105;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#95;&#116;&#125;&#41;&#32;&#92;&#114;&#105;&#103;&#104;&#116;&#97;&#114;&#114;&#111;&#119;&#32;&#091;&#48;&#44;&#32;&#92;&#105;&#110;&#102;&#116;&#121;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="151" style="vertical-align: -5px;"/> that each robot in the fleet uses to assign itself a priority score. Similar to scheduling theory, higher priority robots are more likely to receive human attention. Fleet-DAgger is general enough to model a wide range of IFL algorithms, including IFL adaptations of existing single-robot, single-human IIL algorithms such as <a href="https://arxiv.org/abs/1807.08364" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">EnsembleDAgger</a> and <a href="https://arxiv.org/abs/2109.08273" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ThriftyDAgger</a>. Note, however, that the IFL formalism isn’t limited to Fleet-DAgger: policy learning could be performed with a reinforcement learning algorithm like <a href="https://arxiv.org/abs/1707.06347" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PPO</a>, for instance.</p>
<h2 id="ifl-benchmark-and-experiments">IFL Benchmark and Experiments</h2>
<p>To determine how to best allocate limited human attention to large robot fleets, we need to be able to empirically evaluate and compare different IFL algorithms. To this end, we introduce the <a href="https://github.com/BerkeleyAutomation/ifl_benchmark" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">IFL Benchmark</a>, an open-source Python toolkit available on Github to facilitate the development and standardized evaluation of new IFL algorithms. We extend <a href="https://developer.nvidia.com/isaac-gym" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">NVIDIA Isaac Gym</a>, a highly optimized software library for end-to-end GPU-accelerated robot learning released in 2021, without which the simulation of hundreds or thousands of learning robots would be computationally intractable. Using the IFL Benchmark, we run large-scale simulation experiments with <em>N</em> = 100 robots, <em>M</em> = 10 algorithmic humans, 5 IFL algorithms, and 3 high-dimensional continuous control environments (Figure 1, left).</p>
<p>We also evaluate IFL algorithms in a real-world image-based block pushing task with <em>N</em> = 4 robot arms and <em>M</em> = 2 remote human teleoperators (Figure 1, right). The 4 arms belong to 2 bimanual ABB YuMi robots operating simultaneously in 2 separate labs about 1 kilometer apart, and remote humans in a third physical location perform teleoperation through a keyboard interface when requested. Each robot pushes a cube toward a unique goal position randomly sampled in the workspace; the goals are programmatically generated in the robots’ overhead image observations and automatically resampled when the previous goals are reached. Physical experiment results suggest trends that are approximately consistent with those observed in the benchmark environments.</p>
<h2 id="takeaways-and-future-directions">Takeaways and Future Directions</h2>
<p>To address the gap between the theory and practice of robot fleet learning as well as facilitate future research, we introduce new formalisms, algorithms, and benchmarks for Interactive Fleet Learning. Since IFL does not dictate a specific form or architecture for the shared robot control policy, it can be flexibly synthesized with other promising research directions. For instance, <a href="https://arxiv.org/abs/2303.04137" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">diffusion policies</a>, recently demonstrated to gracefully handle multimodal data, can be used in IFL to allow heterogeneous human supervisor policies. Alternatively, multi-task language-conditioned Transformers like <a href="https://arxiv.org/abs/2212.06817" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RT-1</a> and <a href="https://arxiv.org/abs/2209.05451" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PerAct</a> can be effective “data sponges” that enable the robots in the fleet to perform heterogeneous tasks despite sharing a single policy. The systems aspect of IFL is another compelling research direction: recent developments in cloud and <a href="https://arxiv.org/abs/2205.09778" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">fog robotics</a> enable robot fleets to offload all supervisor allocation, model training, and crowdsourced teleoperation to centralized servers in the cloud with minimal network latency.</p>
<p>While <a href="https://en.wikipedia.org/wiki/Moravec%27s_paradox" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Moravec’s Paradox</a> has so far prevented robotics and embodied AI from fully enjoying the recent spectacular success that Large Language Models (LLMs) like <a href="https://openai.com/research/gpt-4" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">GPT-4</a> have demonstrated, the <a href="http://www.incompleteideas.net/IncIdeas/BitterLesson.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">“bitter lesson”</a> of LLMs is that supervised learning at unprecedented scale is what ultimately leads to the emergent properties we observe. Since we don’t yet have a supply of robot control data nearly as plentiful as all the text and image data on the Internet, the IFL paradigm offers one path forward for scaling up supervised robot learning and deploying robot fleets reliably in today’s world.</p>
<p><em>This post is based on the paper “Fleet-DAgger: Interactive Robot Fleet Learning with Scalable Human Supervision” by Ryan Hoque, Lawrence Chen, Satvik Sharma, Karthik Dharmarajan, Brijen Thananjeyan, Pieter Abbeel, and Ken Goldberg, presented at the Conference on Robot Learning (CoRL) 2022. For more details, see the <a href="https://arxiv.org/abs/2206.14349" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a> on arXiv, <a href="https://www.youtube.com/watch?v=USr_iICRgvk" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">CoRL presentation video</a> on YouTube, open-source <a href="https://github.com/BerkeleyAutomation/ifl_benchmark" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">codebase</a> on Github, <a href="https://twitter.com/ryan_hoque/status/1542932195949432832?s=20" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">high-level summary</a> on Twitter, and <a href="https://sites.google.com/berkeley.edu/fleet-dagger/home" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">project website</a>.</em></p>
<p><em>If you would like to cite this article, please use the following bibtex:</em></p>
<div class="language-plaintext highlighter-rouge">
<div class="highlight">
<pre class="highlight"><code>@article{ifl_blog,
    title={Interactive Fleet Learning},
    author={Hoque, Ryan},
    url={https://bair.berkeley.edu/blog/2023/04/06/ifl/},
    journal={Berkeley Artificial Intelligence Research Blog},
    year={2023} 
}
</code></pre>
</div>
</div>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Fully autonomous real-world reinforcement learning with applications to mobile manipulation</title>
		<link>https://robohub.org/fully-autonomous-real-world-reinforcement-learning-with-applications-to-mobile-manipulation/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 22 Feb 2023 12:59:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://bair.berkeley.edu/blog/2023/01/20/relmm/</guid>

					<description><![CDATA[











Reinforcement learning provides a conceptual framework for autonomous agents to learn from experience, analogously to how one might train a pet with treats. But practical applications of reinforcement learning are often far from natural: i...]]></description>
										<content:encoded><![CDATA[<img decoding="async" src="https://robohub.org/wp-content/uploads/2023/01/sysfig-1024x291.jpg" alt="" width="1024" height="291" class="alignnone size-large wp-image-206660" srcset="https://robohub.org/wp-content/uploads/2023/01/sysfig-1024x291.jpg 1024w, https://robohub.org/wp-content/uploads/2023/01/sysfig-425x121.jpg 425w, https://robohub.org/wp-content/uploads/2023/01/sysfig-768x218.jpg 768w, https://robohub.org/wp-content/uploads/2023/01/sysfig.jpg 1090w" sizes="(max-width: 1024px) 100vw, 1024px" />
<p><strong>By Jędrzej Orbik, Charles Sun, Coline Devin, Glen Berseth</strong></p>
<p>Reinforcement learning provides a conceptual framework for autonomous agents to learn from experience, analogously to how one might train a pet with treats. But practical applications of reinforcement learning are often far from natural: instead of using RL to learn through trial and error by actually attempting the desired task, typical RL applications use a separate (usually simulated) training phase. For example, <a href="https://deepmind.com/research/case-studies/alphago-the-story-so-far" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">AlphaGo</a> did not learn to play Go by competing against thousands of humans, but rather by playing against itself in simulation. While this kind of simulated training is appealing for games where the rules are perfectly known, applying this to real world domains such as robotics can require a range of complex approaches, such as <a href="https://www.youtube.com/watch?v=XUW0cnvqbwM" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">the use of simulated data</a>, or instrumenting real-world environments in various ways to make training feasible <a href="https://bair.berkeley.edu/blog/2020/04/27/ingredients/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">under laboratory conditions</a>. Can we instead devise reinforcement learning systems for robots that allow them to learn directly “on-the-job”, while performing the task that they are required to do? In this blog post, we will discuss ReLMM, a system that we developed that learns to clean up a room directly with a real robot via continual learning.</p>
<p><img decoding="async" src="https://bair.berkeley.edu/static/blog/relmm/image8.gif" width="48%" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/relmm/image12.gif" width="48%" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/relmm/image3.gif" width="48%" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/relmm/image2.gif" width="48%" /><br />
<i>We evaluate our method on different tasks that range in difficulty. The top-left task has uniform white blobs to pickup with no obstacles, while other rooms have objects of diverse shapes and colors, obstacles that increase navigation difficulty and obscure the objects and patterned rugs that make it difficult to see the objects against the ground.</i></p>
<p><span id="more-206407"></span></p>
<p>To enable “on-the-job” training in the real world, the difficulty of collecting more experience is prohibitive. If we can make training in the real world easier, by making the data gathering process more autonomous without requiring human monitoring or intervention, we can further benefit from the simplicity of agents that learn from experience. In this work, we design an “on-the-job” mobile robot training system for cleaning by learning to grasp objects throughout different rooms.</p>
<h2 id="lesson-1-the-benefits-of-modular-policies-for-robots">Lesson 1: The Benefits of Modular Policies for Robots.</h2>
<p>People are not born one day and performing job interviews the next. There are many levels of tasks people learn before they apply for a job as we start with the easier ones and build on them. In ReLMM, we make use of this concept by allowing robots to train common-reusable skills, such as grasping, by first encouraging the robot to prioritize training these skills before learning later skills, such as navigation. Learning in this fashion has two advantages for robotics. The first advantage is that when an agent focuses on learning a skill, it is more efficient at collecting data around the local state distribution for that skill.</p>
<p style="text-align:center">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/relmm/image13.png" width="50%" /><br />

</p>
<p>That is shown in the figure above, where we evaluated the amount of prioritized grasping experience needed to result in efficient mobile manipulation training. The second advantage to a multi-level learning approach is that we can inspect the models trained for different tasks and ask them questions, such as, “can you grasp anything right now” which is helpful for navigation training that we describe next.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/relmm/image14.png" width="50%" /><br />

</p>
<p>Training this multi-level policy was not only more efficient than learning both skills at the same time but it allowed for the grasping controller to inform the navigation policy. Having a model that estimates the uncertainty in its grasp success (<strong>Ours</strong> above) can be used to improve navigation exploration by skipping areas without graspable objects, in contrast to <strong>No Uncertainty Bonus</strong> which does not use this information. The model can also be used to relabel data during training so that in the unlucky case when the grasping model was unsuccessful trying to grasp an object within its reach, the grasping policy can still provide some signal by indicating that an object was there but the grasping policy has not yet learned how to grasp it. Moreover, learning modular models has engineering benefits. Modular training allows for reusing skills that are easier to learn and can enable building intelligent systems one piece at a time. This is beneficial for many reasons, including safety evaluation and understanding.</p>
<h2 id="lesson-2-learning-systems-beat-hand-coded-systems-given-time">Lesson 2: Learning systems beat hand-coded systems, given time</h2>
<p style="text-align:center">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/relmm/image15.png" width="50%" /><br />

</p>
<p>Many robotics tasks that we see today can be solved to varying levels of success using hand-engineered controllers. For our room cleaning task, we designed a hand-engineered controller that locates objects using image clustering and turns towards the nearest detected object at each step. This expertly designed controller performs very well on the visually salient balled socks and takes reasonable paths around the obstacles <strong>but it can not learn an optimal path to collect the objects quickly, and it struggles with visually diverse rooms</strong>. As shown in video 3 below, the scripted policy gets distracted by the white patterned carpet while trying to locate more white objects to grasp.</p>
<p>1) <img decoding="async" src="https://bair.berkeley.edu/static/blog/relmm/image5.gif" width="45%" /><br />
2) <img decoding="async" src="https://bair.berkeley.edu/static/blog/relmm/image6.gif" width="45%" /><br />
3) <img decoding="async" src="https://bair.berkeley.edu/static/blog/relmm/image1.gif" width="45%" /><br />
4) <img decoding="async" src="https://bair.berkeley.edu/static/blog/relmm/image9.png" width="45%" /><br />
<i>We show a comparison between (1) our policy at the beginning of training (2) our policy at the end of training (3) the scripted policy. In (4) we can see the robot&#8217;s performance improve over time, and eventually exceed the scripted policy at quickly collecting the objects in the room.</i>
</p>
<p>Given we can use experts to code this hand-engineered controller, what is the purpose of learning? An important limitation of hand-engineered controllers is that they are tuned for a particular task, for example, grasping white objects. When diverse objects are introduced, which differ in color and shape, the original tuning may no longer be optimal. Rather than requiring further hand-engineering, our learning-based method is able to adapt itself to various tasks by collecting its own experience.</p>
<p>However, the most important lesson is that even if the hand-engineered controller is capable, the learning agent eventually surpasses it given enough time. This learning process is itself autonomous and takes place while the robot is performing its job, making it comparatively inexpensive. This shows the capability of learning agents, which can also be thought of as working out a general way to perform an “expert manual tuning” process for any kind of task. Learning systems have the ability to create the entire control algorithm for the robot, and are not limited to tuning a few parameters in a script. The key step in this work allows these real-world learning systems to autonomously collect the data needed to enable the success of learning methods.</p>
<p><i>This post is based on the paper “Fully Autonomous Real-World Reinforcement Learning with Applications to Mobile Manipulation”, presented at CoRL 2021. You can find more details in <a href="https://arxiv.org/abs/2107.13545" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our paper</a>, on our <a href="https://sites.google.com/view/relmm" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">website</a> and the on the <a href="https://drive.google.com/file/d/1BsqXvxv0ByGIXxGb3zBYBncL9pKaxWuX/view?usp=sharing" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">video</a>. We provide <a href="https://github.com/charlesjsun/ReLMM" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">code</a> to reproduce our experiments. We thank Sergey Levine for his valuable feedback on this blog post.</i></p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Why do Policy Gradient Methods work so well in Cooperative MARL? Evidence from Policy Representation</title>
		<link>https://robohub.org/why-do-policy-gradient-methods-work-so-well-in-cooperative-marl-evidence-from-policy-representation/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sat, 16 Jul 2022 17:00:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://bair.berkeley.edu/blog/2022/07/10/pg-ar/</guid>

					<description><![CDATA[









In cooperative multi-agent reinforcement learning (MARL), due to its on-policy nature, policy gradient (PG) methods are typically believed to be less sample efficient than value decomposition (VD) methods, which are off-policy. However, so...]]></description>
										<content:encoded><![CDATA[<p>In cooperative multi-agent reinforcement learning (MARL), due to its <em>on-policy</em> nature, policy gradient (PG) methods are typically believed to be less sample efficient than value decomposition (VD) methods, which are <em>off-policy</em>. However, some <a href="https://arxiv.org/abs/2103.01955" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recent</a> <a href="https://arxiv.org/abs/2011.09533" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">empirical</a> <a href="https://arxiv.org/abs/2006.07869" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">studies</a> demonstrate that with proper input representation and hyper-parameter tuning, multi-agent PG can achieve <a href="http://bair.berkeley.edu/blog/2021/07/14/mappo/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">surprisingly strong performance</a> compared to off-policy VD methods.</p>
<p><strong>Why could PG methods work so well?</strong> In this post, we will present concrete analysis to show that in certain scenarios, e.g., environments with a highly multi-modal reward landscape, VD can be problematic and lead to undesired outcomes. By contrast, PG methods with individual policies can converge to an optimal policy in these cases. In addition, PG methods with auto-regressive (AR) policies can learn multi-modal policies.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/pg-ar/ar.png" width="80%" /><br />
    <br />
<i><br />
Figure 1: different policy representation for the 4-player permutation game.<br />
</i>
</p>
<p><span id="more-205019"></span></p>
<h2 id="ctde-in-cooperative-marl-vd-and-pg-methods">CTDE in Cooperative MARL: VD and PG methods</h2>
<p>Centralized training and decentralized execution (<a href="https://arxiv.org/abs/1706.02275" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">CTDE</a>) is a popular framework in cooperative MARL. It leverages <em>global</em> information for more effective training while keeping the representation of individual policies for testing. CTDE can be implemented via value decomposition (VD) or policy gradient (PG), leading to two different types of algorithms.</p>
<p>VD methods learn local Q networks and a mixing function that mixes the local Q networks to a global Q function. The mixing function is usually enforced to satisfy the Individual-Global-Max (<a href="https://arxiv.org/abs/1905.05408" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">IGM</a>) principle, which guarantees the optimal joint action can be computed by greedily choosing the optimal action locally for each agent.</p>
<p>By contrast, PG methods directly apply policy gradient to learn an individual policy and a centralized value function for each agent. The value function takes as its input the global state (e.g., <a href="https://arxiv.org/abs/2103.01955" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">MAPPO</a>) or the concatenation of all the local observations (e.g., <a href="https://arxiv.org/abs/1706.02275" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">MADDPG</a>), for an accurate global value estimate.</p>
<h2 id="the-permutation-game-a-simple-counterexample-where-vd-fails">The permutation game: a simple counterexample where VD fails</h2>
<p>We start our analysis by considering a stateless cooperative game, namely the permutation game. In an <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-5793832f979c2268e3694c246d53b1bb_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#78;" title="Rendered by QuickLaTeX.com" height="12" width="16" style="vertical-align: 0px;"/>-player permutation game, each agent can output <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-5793832f979c2268e3694c246d53b1bb_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#78;" title="Rendered by QuickLaTeX.com" height="12" width="16" style="vertical-align: 0px;"/> actions <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-13fdd49e0bb70eb52b23f39172621e09_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#123;&#32;&#49;&#44;&#92;&#108;&#100;&#111;&#116;&#115;&#44;&#32;&#78;&#32;&#125;" title="Rendered by QuickLaTeX.com" height="16" width="63" style="vertical-align: -4px;"/>. Agents receive <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-d4932b1a489a055f2908670bac049524_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#43;&#49;" title="Rendered by QuickLaTeX.com" height="14" width="22" style="vertical-align: -2px;"/> reward  if their actions are mutually different, i.e., the joint action is a permutation over <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-b784246062f3bc4dec1f6939e7320995_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#49;&#44;&#32;&#92;&#108;&#100;&#111;&#116;&#115;&#44;&#32;&#78;" title="Rendered by QuickLaTeX.com" height="16" width="63" style="vertical-align: -4px;"/>; otherwise, they receive <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-a5e437be25f29374d30f66cd46adf81c_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#48;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: 0px;"/> reward. Note that there are <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ace60da110c06c63e44172f3a6dffe04_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#78;&#33;" title="Rendered by QuickLaTeX.com" height="13" width="20" style="vertical-align: 0px;"/> symmetric optimal strategies in this game.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/pg-ar/permutation_game.png" width="70%" /><br />
    <br />
<i><br />
Figure 2: the 4-player permutation game.<br />
</i>
</p>
<p>Let us focus on the 2-player permutation game for our discussion. In this setting, if we apply VD to the game, the global Q-value will factorize to</p>
<p class="ql-center-displayed-equation" style="line-height: 22px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-7cbe906c4064e7c74f4789099232a35f_l3.png" height="22" width="275" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#97;&#94;&#49;&#44;&#97;&#94;&#50;&#41;&#61;&#102;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#109;&#105;&#120;&#125;&#40;&#81;&#95;&#49;&#40;&#97;&#94;&#49;&#41;&#44;&#81;&#95;&#50;&#40;&#97;&#94;&#50;&#41;&#41;&#44;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>where <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-6b6b12b50fe43a209ecce557539ee185_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#95;&#49;" title="Rendered by QuickLaTeX.com" height="16" width="20" style="vertical-align: -4px;"/> and <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-7a8f7f7bc05736504761c873bfd99aa1_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#95;&#50;" title="Rendered by QuickLaTeX.com" height="16" width="21" style="vertical-align: -4px;"/> are local Q-functions, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-a0e07607931a066cb71e9478aa5daea8_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;" title="Rendered by QuickLaTeX.com" height="16" width="31" style="vertical-align: -4px;"/> is the global Q-function, and <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-c097142d659c1a925ea1d35cfa7a00bb_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#102;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#109;&#105;&#120;&#125;" title="Rendered by QuickLaTeX.com" height="16" width="32" style="vertical-align: -4px;"/> is the mixing function that, as required by VD methods, satisfies the IGM principle.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/pg-ar/vd_pg.png" width="90%" /><br />
    <br />
    <i><br />
Figure 3: high-level intuition on why VD fails in the 2-player permutation game.<br />
    </i>
</p>
<p>We formally prove that VD cannot represent the payoff of the 2-player permutation game by contradiction. If VD methods were able to represent the payoff, we would have</p>
<p class="ql-center-displayed-equation" style="line-height: 19px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-fbacf98e0b8a45e370a60629a93885d4_l3.png" height="19" width="503" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#49;&#44;&#32;&#50;&#41;&#61;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#50;&#44;&#49;&#41;&#61;&#49;&#32;&#92;&#113;&#113;&#117;&#97;&#100;&#32;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#97;&#110;&#100;&#125;&#32;&#92;&#113;&#113;&#117;&#97;&#100;&#32;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#49;&#44;&#32;&#49;&#41;&#61;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#50;&#44;&#50;&#41;&#61;&#48;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>However, if either of these two agents have different local Q values, e.g. <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-9eb72ff54a6609accc2f617796dd96e5_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#95;&#49;&#40;&#49;&#41;&#62;&#32;&#81;&#95;&#49;&#40;&#50;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="112" style="vertical-align: -5px;"/>, then according to the IGM principle, we must have</p>
<p class="ql-center-displayed-equation" style="line-height: 29px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3720cf145461d912bdf0414d5f358785_l3.png" height="29" width="570" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#49;&#61;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#49;&#44;&#50;&#41;&#61;&#92;&#97;&#114;&#103;&#92;&#109;&#97;&#120;&#95;&#123;&#97;&#94;&#50;&#125;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#49;&#44;&#97;&#94;&#50;&#41;&#62;&#92;&#97;&#114;&#103;&#92;&#109;&#97;&#120;&#95;&#123;&#97;&#94;&#50;&#125;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#50;&#44;&#97;&#94;&#50;&#41;&#61;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#50;&#44;&#49;&#41;&#61;&#49;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>Otherwise, if <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3c7fcbe3433fbb756f4c7fdad6d5424b_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#95;&#49;&#40;&#49;&#41;&#61;&#81;&#95;&#49;&#40;&#50;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="112" style="vertical-align: -5px;"/> and <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-c8eb741acebd31f603795eb0fb944811_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#95;&#50;&#40;&#49;&#41;&#61;&#81;&#95;&#50;&#40;&#50;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="112" style="vertical-align: -5px;"/>, then</p>
<p class="ql-center-displayed-equation" style="line-height: 19px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-dcb344ca51d74548f2f54102716919a9_l3.png" height="19" width="363" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#49;&#44;&#32;&#49;&#41;&#61;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#50;&#44;&#50;&#41;&#61;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#49;&#44;&#32;&#50;&#41;&#61;&#81;&#95;&#92;&#116;&#101;&#120;&#116;&#114;&#109;&#123;&#116;&#111;&#116;&#125;&#40;&#50;&#44;&#49;&#41;&#46;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>As a result, value decomposition cannot represent the payoff matrix of the 2-player permutation game.</p>
<p>What about PG methods? Individual policies can indeed represent an optimal policy for the permutation game. Moreover, stochastic gradient descent can guarantee PG to converge to one of these optima <a href="https://arxiv.org/abs/1802.06175" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">under mild assumptions</a>. This suggests that, even though PG methods are less popular in MARL compared with VD methods, they can be preferable in certain cases that are common in real-world applications, e.g., games with multiple strategy modalities.</p>
<p>We also remark that in the permutation game, in order to represent an optimal joint policy, each agent must choose distinct actions. <strong>Consequently, a successful implementation of PG must ensure that the policies are agent-specific.</strong> This can be done by using either individual policies with unshared parameters (referred to as PG-Ind in our paper), or an agent-ID conditioned policy (<a href="http://bair.berkeley.edu/blog/2021/07/14/mappo/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PG-ID</a>).</p>
<h2 id="pg-outperform-best-vd-methods-on-popular-marl-testbeds">PG outperform best VD methods on popular MARL testbeds</h2>
<p>Going beyond the simple illustrative example of the permutation game, we extend our study to popular and more realistic MARL benchmarks. In addition to StarCraft Multi-Agent Challenge (<a href="https://github.com/oxwhirl/smac" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SMAC</a>), where the effectiveness of PG and agent-conditioned policy input <a href="http://bair.berkeley.edu/blog/2021/07/14/mappo/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">has been verified</a>, we show new results in Google Research Football (<a href="https://github.com/google-research/football" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">GRF</a>) and multi-player <a href="https://github.com/deepmind/hanabi-learning-environment" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Hanabi Challenge</a>.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/pg-ar/football.png" width="48%" /><br />
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/pg-ar/hanabi.png" width="45%" /><br />
    <br />
<i><br />
Figure 4: (top) winning rates of PG methods on GRF; (bottom) best and average evaluation scores on Hanabi-Full.<br />
</i>
</p>
<p>In GRF, PG methods outperform the state-of-the-art VD baseline (<a href="https://arxiv.org/abs/2106.02195" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">CDS</a>) in 5 scenarios. Interestingly, we also notice that individual policies (PG-Ind) without parameter sharing achieve comparable, sometimes even higher winning rates, compared to agent-specific policies (PG-ID) in all 5 scenarios. We evaluate PG-ID in the full-scale Hanabi game with varying numbers of players (2-5 players) and compare them to <a href="https://arxiv.org/abs/1912.02288" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SAD</a>, a strong off-policy Q-learning variant in Hanabi, and Value Decomposition Networks (<a href="https://arxiv.org/abs/1706.05296" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">VDN</a>). As demonstrated in the above table, PG-ID is able to produce results comparable to or better than the best and average rewards achieved by SAD and VDN with varying numbers of players using the same number of environment steps.</p>
<h2 id="beyond-higher-rewards-learning-multi-modal-behavior-via-auto-regressive-policy-modeling">Beyond higher rewards: learning multi-modal behavior via auto-regressive policy modeling</h2>
<p>Besides learning higher rewards, we also study how to learn multi-modal policies in cooperative MARL. Let’s go back to the permutation game. Although we have proved that PG can effectively learn an optimal policy, the strategy mode that it finally reaches can highly depend on the policy initialization. Thus, a natural question will be:</p>
<p style="text-align:center;">
    <i><br />
Can we learn a single policy that can cover all the optimal modes?<br />
    </i>
</p>
<p>In the decentralized PG formulation, the factorized representation of a joint policy can only represent one particular mode. Therefore, we propose an enhanced way to parameterize the policies for stronger expressiveness — the auto-regressive (AR) policies.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/pg-ar/permutation_ar.gif" width="80%" /><br />
<br />
<i><br />
Figure 5: comparison between individual policies (PG) and auto-regressive  policies (AR) in the 4-player permutation game.<br />
</i>
</p>
<p>Formally, we factorize the joint policy of <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-b170995d512c659d8668b4e42e1fef6b_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#110;" title="Rendered by QuickLaTeX.com" height="8" width="11" style="vertical-align: 0px;"/> agents into the form of</p>
<p class="ql-center-displayed-equation" style="line-height: 49px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-4996f8d6a765353aca34c32c83c9dfe1_l3.png" height="49" width="298" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#112;&#105;&#40;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#97;&#125;&#32;&#92;&#109;&#105;&#100;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#102;&#123;&#111;&#125;&#41;&#32;&#92;&#97;&#112;&#112;&#114;&#111;&#120;&#32;&#92;&#112;&#114;&#111;&#100;&#95;&#123;&#105;&#61;&#49;&#125;&#94;&#110;&#32;&#92;&#112;&#105;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#94;&#123;&#105;&#125;&#125;&#32;&#92;&#108;&#101;&#102;&#116;&#40;&#32;&#97;&#94;&#123;&#105;&#125;&#92;&#109;&#105;&#100;&#32;&#111;&#94;&#123;&#105;&#125;&#44;&#97;&#94;&#123;&#49;&#125;&#44;&#92;&#108;&#100;&#111;&#116;&#115;&#44;&#97;&#94;&#123;&#105;&#45;&#49;&#125;&#32;&#92;&#114;&#105;&#103;&#104;&#116;&#41;&#44;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>where the action produced by agent <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-695d9d59bd04859c6c99e7feb11daab6_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#105;" title="Rendered by QuickLaTeX.com" height="12" width="6" style="vertical-align: 0px;"/> depends on its own observation <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-040cf8536524b550ec19e15c2eeffcc7_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#111;&#95;&#105;" title="Rendered by QuickLaTeX.com" height="11" width="14" style="vertical-align: -3px;"/> and all the actions from previous agents <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ebdb2ba9f9bc68dcb1efa023c87d2699_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#49;&#44;&#92;&#100;&#111;&#116;&#115;&#44;&#105;&#45;&#49;" title="Rendered by QuickLaTeX.com" height="16" width="83" style="vertical-align: -4px;"/>. The auto-regressive factorization can represent <em>any</em> joint policy in a centralized MDP. The <em>only</em> modification to each agent’s policy is the input dimension, which is slightly enlarged by including previous actions; and the output dimension of each agent’s policy remains unchanged.</p>
<p>With such a minimal parameterization overhead, AR policy substantially improves the representation power of PG methods. We remark that PG with AR policy (PG-AR) can simultaneously represent all optimal policy modes in the permutation game.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/pg-ar/heatmap.png" width="70%" /><br />
    <br />
<i><br />
Figure: the heatmaps of actions for policies learned by PG-Ind (left) and PG-AR (middle), and the heatmap for rewards (right); while PG-Ind only converge to a specific mode in the 4-player permutation game, PG-AR successfully discovers all the optimal modes.<br />
</i>
</p>
<p>In more complex environments, including SMAC and GRF, PG-AR can learn interesting emergent behaviors that require strong intra-agent coordination that may never be learned by PG-Ind.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/pg-ar/2m1z.gif" width="45%" /><br />
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/pg-ar/3v1.gif" width="45%" /><br />
    <br />
<i><br />
Figure 6: (top) emergent behavior induced by PG-AR in SMAC and GRF. On the 2m_vs_1z map of SMAC, the marines keep standing and attack alternately while ensuring there is only one attacking marine at each timestep; (bottom) in the academy_3_vs_1_with_keeper scenario of GRF, agents learn a &#8220;Tiki-Taka&#8221; style behavior: each player keeps passing the ball to their teammates.<br />
</i>
</p>
<h2 id="discussions-and-takeaways">Discussions and Takeaways</h2>
<p>In this post, we provide a concrete analysis of VD and PG methods in cooperative MARL. First, we reveal the limitation on the expressiveness of popular VD methods, showing that they could not represent optimal policies even in a simple permutation game. By contrast, we show that PG methods are provably more expressive. We empirically verify the expressiveness advantage of PG on popular MARL testbeds, including SMAC, GRF, and Hanabi Challenge. We hope the insights from this work could benefit the community towards more general and more powerful cooperative MARL algorithms in the future.</p>
<hr />
<p><em>This post is based on our paper in joint with Zelai Xu: Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement Learning (<a href="https://arxiv.org/abs/2206.07505" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a>, <a href="https://sites.google.com/view/revisiting-marl" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">website</a>).</em></p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Designing societally beneficial Reinforcement Learning (RL) systems</title>
		<link>https://robohub.org/designing-societally-beneficial-reinforcement-learning-rl-systems/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sun, 15 May 2022 10:00:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<category><![CDATA[research]]></category>
		<guid isPermaLink="false">http://bair.berkeley.edu/blog/2022/04/29/reward-reports/</guid>

					<description><![CDATA[












Deep reinforcement learning (DRL) is transitioning from a research field focused on game playing to a technology with real-world applications. Notable examples include DeepMind’s work on controlling a nuclear reactor or on improving Yout...]]></description>
										<content:encoded><![CDATA[<p><strong>By Nathan Lambert, Aaron Snoswell, Sarah Dean, Thomas Krendl Gilbert, and Tom Zick</strong></p>
<p>Deep reinforcement learning (DRL) is transitioning from a research field focused on game playing to a technology with real-world applications. Notable examples include DeepMind’s work on <a href="https://www.nature.com/articles/s41586-021-04301-9" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">controlling a nuclear reactor</a> or on improving <a href="https://arxiv.org/abs/2202.06626" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Youtube video compression</a>, or Tesla <a href="https://www.youtube.com/watch?v=j0z4FweCy4M&amp;t=4802s" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">attempting to use a method inspired by MuZero</a> for autonomous vehicle behavior planning. But the exciting potential for real world applications of RL should also come with a healthy dose of caution &#8211; for example RL policies are well known to be vulnerable to <a href="https://robotic.substack.com/p/rl-exploitation?s=w" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">exploitation</a>, and methods for safe and <a href="https://bair.berkeley.edu/blog/2021/03/09/maxent-robust-rl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robust policy development</a> are an active area of research.</p>
<p>At the same time as the emergence of powerful RL systems in the real world, the public and researchers are expressing an increased appetite for fair, aligned, and safe machine learning systems. The focus of these research efforts to date has been to account for shortcomings of datasets or supervised learning practices that can harm individuals. However the unique ability of RL systems to leverage temporal feedback in learning complicates the types of risks and safety concerns that can arise.</p>
<p>This post expands on our recent <a href="https://cltc.berkeley.edu/2022/02/08/reward-reports/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">whitepaper</a> and <a href="https://arxiv.org/abs/2204.10817" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">research paper</a>, where we aim to illustrate the different modalities harms can take when augmented with the temporal axis of RL. To combat these novel societal risks, we also propose a new kind of documentation for dynamic Machine Learning systems which aims to assess and monitor these risks both before and after deployment.</p>
<p><span id="more-204301"></span></p>
<h2 id="whats-special-about-rl-a-taxonomy-of-feedback">What’s Special About RL? A Taxonomy of Feedback</h2>
<p>Reinforcement learning systems are often spotlighted for their ability to act in an environment, rather than passively make predictions. Other supervised machine learning systems, such as computer vision, consume data and return a prediction that can be used by some decision making rule. In contrast, the appeal of RL is in its ability to not only (a) directly model the impact of actions, but also to (b) improve policy performance automatically. These key properties of acting upon an environment, and learning within that environment can be understood as by considering the different types of feedback that come into play when an RL agent acts within an environment. We classify these feedback forms in a taxonomy of (1) Control, (2) Behavioral, and (3) Exogenous feedback. The first two notions of feedback, Control and Behavioral, are directly within the formal mathematical definition of an RL agent while Exogenous feedback is induced as the agent interacts with the broader world.</p>
<h3 id="1-control-feedback">1. Control Feedback</h3>
<p>First is control feedback &#8211; in the control systems engineering sense &#8211; where the action taken depends on the current measurements of the state of the system. RL agents choose actions based on an observed state according to a policy, which generates environmental feedback. For example, a thermostat turns on a furnace according to the current temperature measurement. Control feedback gives an agent the ability to react to unforeseen events (e.g. a sudden snap of cold weather) autonomously.</p>
<div id="attachment_204435" style="width: 1034px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-204435" src="https://robohub.org/wp-content/uploads/2022/04/fb-control-1024x546.png" alt="" width="1024" height="546" class="size-large wp-image-204435" srcset="https://robohub.org/wp-content/uploads/2022/04/fb-control-1024x546.png 1024w, https://robohub.org/wp-content/uploads/2022/04/fb-control-425x227.png 425w, https://robohub.org/wp-content/uploads/2022/04/fb-control-768x409.png 768w, https://robohub.org/wp-content/uploads/2022/04/fb-control-1536x819.png 1536w, https://robohub.org/wp-content/uploads/2022/04/fb-control.png 1842w" sizes="(max-width: 1024px) 100vw, 1024px" /><p id="caption-attachment-204435" class="wp-caption-text">Figure 1: Control Feedback.</p></div>
<h3 id="2-behavioral-feedback">2. Behavioral Feedback</h3>
<p>Next in our taxonomy of RL feedback is ‘behavioral feedback’: the trial and error learning that enables an agent to improve its policy through interaction with the environment. This could be considered the defining feature of RL, as compared to e.g. ‘classical’ control theory. Policies in RL can be defined by a set of parameters that determine the actions the agent takes in the future. Because these parameters are updated through behavioral feedback, these are actually a reflection of the data collected from executions of past policy versions. RL agents are not fully ‘memoryless’ in this respect–the current policy depends on stored experience, and impacts newly collected data, which in turn impacts future versions of the agent. To continue the thermostat example &#8211; a ‘smart home’ thermostat might analyze historical temperature measurements and adapt its control parameters in accordance with seasonal shifts in temperature, for instance to have a more aggressive control scheme during winter months.</p>
<div id="attachment_204436" style="width: 976px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-204436" src="https://robohub.org/wp-content/uploads/2022/04/fb-behavioral.png" alt="" width="966" height="740" class="size-full wp-image-204436" srcset="https://robohub.org/wp-content/uploads/2022/04/fb-behavioral.png 966w, https://robohub.org/wp-content/uploads/2022/04/fb-behavioral-425x326.png 425w, https://robohub.org/wp-content/uploads/2022/04/fb-behavioral-768x588.png 768w" sizes="(max-width: 966px) 100vw, 966px" /><p id="caption-attachment-204436" class="wp-caption-text">Figure 2: Behavioral Feedback.</p></div>
<h3 id="3-exogenous-feedback">3. Exogenous Feedback</h3>
<p>Finally, we can consider a third form of feedback external to the specified RL environment, which we call Exogenous (or ‘exo’) feedback. While RL benchmarking tasks may be static environments, every action in the real world impacts the dynamics of both the target deployment environment, as well as adjacent environments. For example, a news recommendation system that is optimized for clickthrough may change the way editors write headlines towards attention-grabbing  clickbait. In this RL formulation, the set of articles to be recommended would be considered part of the environment and expected to remain static, but exposure incentives cause a shift over time.</p>
<p>To continue the thermostat example, as a ‘smart thermostat’ continues to adapt its behavior over time, the behavior of other adjacent systems in a household might change in response &#8211; for instance other appliances might consume more electricity due to increased heat levels, which could impact electricity costs. Household occupants might also change their clothing and behavior patterns due to different temperature profiles during the day. In turn, these secondary effects could also influence the temperature which the thermostat monitors, leading to a longer timescale feedback loop.</p>
<p>Negative costs of these external effects will not be specified in the agent-centric reward function, leaving these external environments to be manipulated or exploited. Exo-feedback is by definition difficult for a designer to predict. Instead, we propose that it should be addressed by documenting the evolution of the agent, the targeted environment, and adjacent environments.</p>
<div id="attachment_204437" style="width: 1034px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-204437" src="https://robohub.org/wp-content/uploads/2022/04/fb-exo-1024x823.png" alt="" width="1024" height="823" class="size-large wp-image-204437" srcset="https://robohub.org/wp-content/uploads/2022/04/fb-exo-1024x823.png 1024w, https://robohub.org/wp-content/uploads/2022/04/fb-exo-425x341.png 425w, https://robohub.org/wp-content/uploads/2022/04/fb-exo-768x617.png 768w, https://robohub.org/wp-content/uploads/2022/04/fb-exo.png 1414w" sizes="(max-width: 1024px) 100vw, 1024px" /><p id="caption-attachment-204437" class="wp-caption-text">Figure 3: Exogenous (exo) Feedback.</p></div>
<h2 id="how-can-rl-systems-fail">How can RL systems fail?</h2>
<p>Let’s consider how two key properties can lead to failure modes specific to RL systems: direct action selection (via control feedback) and autonomous data collection (via behavioral feedback).</p>
<p>First is decision-time safety. One current practice in RL research to create safe decisions is to augment the agent’s reward function with a penalty term for certain harmful or undesirable states and actions. For example, in a robotics domain we might penalize certain actions (such as extremely large torques) or state-action tuples (such as carrying a glass of water over sensitive equipment). However it is difficult to anticipate where on a pathway an agent may encounter a crucial action, such that failure would result in an unsafe event. This aspect of how reward functions interact with optimizers is especially problematic for deep learning systems, where numerical guarantees are challenging.</p>
<div id="attachment_204438" style="width: 1034px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-204438" src="https://robohub.org/wp-content/uploads/2022/04/decision-1024x479.png" alt="" width="1024" height="479" class="size-large wp-image-204438" srcset="https://robohub.org/wp-content/uploads/2022/04/decision-1024x479.png 1024w, https://robohub.org/wp-content/uploads/2022/04/decision-425x199.png 425w, https://robohub.org/wp-content/uploads/2022/04/decision-768x359.png 768w, https://robohub.org/wp-content/uploads/2022/04/decision.png 1315w" sizes="(max-width: 1024px) 100vw, 1024px" /><p id="caption-attachment-204438" class="wp-caption-text">Figure 4: Decision time failure illustration.</p></div>
<p>As an RL agent collects new data and the policy adapts, there is a complex interplay between current parameters, stored data, and the environment that governs evolution of the system. Changing any one of these three sources of information will change the future behavior of the agent, and moreover these three components are deeply intertwined. This uncertainty makes it difficult to back out the cause of failures or successes.</p>
<p>In domains where many behaviors can possibly be expressed, the RL specification leaves a lot of factors constraining behavior unsaid. For a robot learning locomotion over an uneven environment, it would be useful to know what signals in the system indicate it will learn to find an easier route rather than a more complex gait. In complex situations with less well-defined reward functions, these intended or unintended behaviors will encompass a much broader range of capabilities, which may or may not have been accounted for by the designer.</p>
<div id="attachment_204439" style="width: 1034px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-204439" src="https://robohub.org/wp-content/uploads/2022/04/behavior-1024x743.png" alt="" width="1024" height="743" class="size-large wp-image-204439" srcset="https://robohub.org/wp-content/uploads/2022/04/behavior-1024x743.png 1024w, https://robohub.org/wp-content/uploads/2022/04/behavior-425x309.png 425w, https://robohub.org/wp-content/uploads/2022/04/behavior-768x558.png 768w, https://robohub.org/wp-content/uploads/2022/04/behavior.png 1350w" sizes="(max-width: 1024px) 100vw, 1024px" /><p id="caption-attachment-204439" class="wp-caption-text">Figure 5: Behavior estimation failure illustration.</p></div>
<p>While these failure modes are closely related to control and behavioral feedback, Exo-feedback does not map as clearly to one type of error and introduces risks that do not fit into simple categories. Understanding exo-feedback requires that stakeholders in the broader communities (machine learning, application domains, sociology, etc.) work together on real world RL deployments.</p>
<h2 id="risks-with-real-world-rl">Risks with real-world RL</h2>
<p>Here, we discuss four types of design choices an RL designer must make, and how these choices can have an impact upon the socio-technical failures that an agent might exhibit once deployed.</p>
<h3 id="scoping-the-horizon">Scoping the Horizon</h3>
<p>Determining the timescale on which aRL agent can plan impacts the possible and actual behavior of that agent. In the lab, it may be common to tune the horizon length until the desired behavior is achieved. But in real world systems, optimizations will externalize costs depending on the defined horizon. For example, an RL agent controlling an autonomous vehicle will have very different goals and behaviors if the task is to stay in a lane,  navigate a contested intersection, or route across a city to a destination. This is true even if the objective (e.g. “minimize travel time”) remains the same.</p>
<div id="attachment_204440" style="width: 1034px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-204440" src="https://robohub.org/wp-content/uploads/2022/04/horizon-1024x398.png" alt="" width="1024" height="398" class="size-large wp-image-204440" srcset="https://robohub.org/wp-content/uploads/2022/04/horizon-1024x398.png 1024w, https://robohub.org/wp-content/uploads/2022/04/horizon-425x165.png 425w, https://robohub.org/wp-content/uploads/2022/04/horizon-768x299.png 768w, https://robohub.org/wp-content/uploads/2022/04/horizon-1536x598.png 1536w, https://robohub.org/wp-content/uploads/2022/04/horizon.png 1892w" sizes="(max-width: 1024px) 100vw, 1024px" /><p id="caption-attachment-204440" class="wp-caption-text">Figure 6: Scoping the horizon example with an autonomous vehicle.</p></div>
<h3 id="defining-rewards">Defining Rewards</h3>
<p>A second design choice is that of actually specifying the reward function to be maximized. This immediately raises the well-known risk of RL systems, reward hacking, where the designer and agent negotiate behaviors based on specified reward functions. In a deployed RL system, this often results in unexpected exploitative behavior – from <a href="https://openai.com/blog/faulty-reward-functions/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">bizarre video game agents</a> to <a href="https://bair.berkeley.edu/blog/2021/04/19/mbrl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">causing errors in robotics simulators</a>. For example, if an agent is presented with the problem of navigating a maze to reach the far side, a mis-specified reward might result in the agent avoiding the task entirely to minimize the time taken.</p>
<div id="attachment_204441" style="width: 1034px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-204441" src="https://robohub.org/wp-content/uploads/2022/04/reward-shaping-1024x422.png" alt="" width="1024" height="422" class="size-large wp-image-204441" srcset="https://robohub.org/wp-content/uploads/2022/04/reward-shaping-1024x422.png 1024w, https://robohub.org/wp-content/uploads/2022/04/reward-shaping-425x175.png 425w, https://robohub.org/wp-content/uploads/2022/04/reward-shaping-768x317.png 768w, https://robohub.org/wp-content/uploads/2022/04/reward-shaping-1536x633.png 1536w, https://robohub.org/wp-content/uploads/2022/04/reward-shaping.png 1722w" sizes="(max-width: 1024px) 100vw, 1024px" /><p id="caption-attachment-204441" class="wp-caption-text">Figure 7: Defining rewards example with maze navigation.</p></div>
<h3 id="pruning-information">Pruning Information</h3>
<p>A common practice in RL research is to redefine the environment to fit one’s needs – RL designers make numerous explicit and implicit assumptions to model tasks in a way that makes them amenable to virtual RL agents. In highly structured domains, such as video games, this can be rather benign.However, in the real world redefining the environment amounts to changing the ways information can flow between the world and the RL agent. This can dramatically change the meaning of the reward function and offload risk to external systems. For example, an autonomous vehicle with sensors focused only on the road surface shifts the burden from AV designers to pedestrians. In this case, the designer is pruning out information about the surrounding environment that is actually crucial to robustly safe integration within society.</p>
<div id="attachment_204443" style="width: 1034px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-204443" src="https://robohub.org/wp-content/uploads/2022/04/info-shaping-1024x581.png" alt="" width="1024" height="581" class="size-large wp-image-204443" srcset="https://robohub.org/wp-content/uploads/2022/04/info-shaping-1024x581.png 1024w, https://robohub.org/wp-content/uploads/2022/04/info-shaping-425x241.png 425w, https://robohub.org/wp-content/uploads/2022/04/info-shaping-768x436.png 768w, https://robohub.org/wp-content/uploads/2022/04/info-shaping-1536x871.png 1536w, https://robohub.org/wp-content/uploads/2022/04/info-shaping.png 1728w" sizes="(max-width: 1024px) 100vw, 1024px" /><p id="caption-attachment-204443" class="wp-caption-text">Figure 8: Information shaping example with an autonomous vehicle.</p></div>
<h3 id="training-multiple-agents">Training Multiple Agents</h3>
<p>There is growing interest in the problem of <a href="https://bair.berkeley.edu/blog/2021/07/14/mappo/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">multi-agent RL</a>, but as an emerging research area, little is known about how learning systems interact within dynamic environments. When the relative concentration of autonomous agents increases within an environment, the terms these agents optimize for can actually re-wire norms and values encoded in that specific application domain. An example would be the changes in behavior that will come if the majority of vehicles are autonomous and communicating (or not) with each other. In this case, if the agents have autonomy to optimize toward a goal of minimizing transit time (for example), they could crowd out the remaining human drivers and heavily disrupt accepted societal norms of transit.</p>
<div id="attachment_204442" style="width: 1034px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-204442" src="https://robohub.org/wp-content/uploads/2022/04/multi-agent-1024x693.png" alt="" width="1024" height="693" class="size-large wp-image-204442" srcset="https://robohub.org/wp-content/uploads/2022/04/multi-agent-1024x693.png 1024w, https://robohub.org/wp-content/uploads/2022/04/multi-agent-425x287.png 425w, https://robohub.org/wp-content/uploads/2022/04/multi-agent-768x519.png 768w, https://robohub.org/wp-content/uploads/2022/04/multi-agent-1536x1039.png 1536w, https://robohub.org/wp-content/uploads/2022/04/multi-agent.png 1854w" sizes="(max-width: 1024px) 100vw, 1024px" /><p id="caption-attachment-204442" class="wp-caption-text">Figure 9: The risks of multi-agency example on autonomous vehicles.</p></div>
<h2 id="making-sense-of-applied-rl-reward-reporting">Making sense of applied RL: Reward Reporting</h2>
<p>In our recent <a href="https://cltc.berkeley.edu/2022/02/08/reward-reports/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">whitepaper</a> and <a href="https://arxiv.org/abs/2204.10817" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">research paper</a>, we proposed <a href="https://rewardreports.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Reward Reports</a>, a new form of ML documentation that foregrounds the societal risks posed by sequential data-driven optimization systems, whether explicitly constructed as an RL agent or <a href="https://robotic.substack.com/p/ml-becomes-rl?s=w" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">implicitly construed</a> via data-driven optimization and feedback. Building on proposals to document datasets and models, we focus on reward functions: the objective that guides optimization decisions in feedback-laden systems. Reward Reports comprise questions that highlight the promises and risks entailed in defining what is being optimized in an AI system, and are intended as living documents that dissolve the distinction between ex-ante (design) specification and ex-post (after the fact) harm. As a result, Reward Reports provide a framework for ongoing deliberation and accountability before and after a system is deployed.</p>
<p>Our proposed template for a Reward Reports consists of several sections, arranged to help the reporter themselves understand and document the system. A Reward Report begins with (1) system details that contain the information context for deploying the model. From there, the report documents (2) the optimization intent, which questions the goals of the system and why RL or ML may be a useful tool. The designer then documents (3) how the system may affect different stakeholders in the institutional interface. The next two sections contain technical details on (4) the system implementation and (5) evaluation. Reward reports conclude with (6) plans for system maintenance as additional system dynamics are uncovered.</p>
<p>The most important feature of a Reward Report is that it allows documentation to evolve over time, in step with the temporal evolution of an online, deployed RL system! This is most evident in the change-log, which is we locate at the end of our Reward Report template:</p>
<div id="attachment_204444" style="width: 537px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-204444" src="https://robohub.org/wp-content/uploads/2022/04/rr-contents-527x1024.png" alt="" width="527" height="1024" class="size-large wp-image-204444" srcset="https://robohub.org/wp-content/uploads/2022/04/rr-contents-527x1024.png 527w, https://robohub.org/wp-content/uploads/2022/04/rr-contents-219x425.png 219w, https://robohub.org/wp-content/uploads/2022/04/rr-contents.png 593w" sizes="(max-width: 527px) 100vw, 527px" /><p id="caption-attachment-204444" class="wp-caption-text">Figure 10: Reward Reports contents.</p></div>
<h3 id="what-would-this-look-like-in-practice">What would this look like in practice?</h3>
<p>As part of our research, we have developed a reward report <a href="https://github.com/RewardReports/reward-reports" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">LaTeX template, as well as several example reward reports</a> that aim to illustrate the kinds of issues that could be managed by this form of documentation. These examples include the temporal evolution of the MovieLens recommender system, the DeepMind MuZero game playing system, and a hypothetical deployment of an RL autonomous vehicle policy for managing merging traffic, based on the <a href="https://flow-project.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project Flow simulator</a>.</p>
<p>However, these are just examples that we hope will serve to inspire the RL community–as more RL systems are deployed in real-world applications, we hope the research community will build on our ideas for Reward Reports and refine the specific content that should be included. To this end, we hope that you will join us at our (un)-workshop.</p>
<h3 id="work-with-us-on-reward-reports-an-unworkshop">Work with us on Reward Reports: An (Un)Workshop!</h3>
<p>We are hosting an “un-workshop” at the upcoming conference on Reinforcement Learning and Decision Making (<a href="https://rldm.org/rldm-2022-workshops/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RLDM</a>) on June 11th from 1:00-5:00pm EST at Brown University, Providence, RI. We call this an un-workshop because we are looking for the attendees to help create the content! We will provide templates, ideas, and discussion as our attendees build out example reports. We are excited to develop the ideas behind Reward Reports with real-world practitioners and cutting-edge researchers.</p>
<p>For more information on the workshop, visit the <a href="https://rewardreports.github.io/workshop.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">website</a> or contact the organizers at <a href="mailto:geese-org@lists.berkeley.edu">geese-org@lists.berkeley.edu</a>.</p>
<p>This post is based on the following papers:</p>
<ul>
<li><a href="https://cltc.berkeley.edu/2022/02/08/reward-reports/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Choices, Risks, and Reward Reports: Charting Public Policy for Reinforcement Learning Systems</a> by Thomas Krendl Gilbert, Sarah Dean, Tom Zick, Nathan Lambert. Center for Long Term Cybersecurity Whitepaper Series 2022.</li>
<li><a href="https://arxiv.org/abs/2204.10817" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Reward Reports for Reinforcement Learning</a> by Thomas Krendl Gilbert, Sarah Dean, Nathan Lambert, Tom Zick and Aaron Snoswell. ArXiv Preprint 2022.</li>
</ul>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>RECON: Learning to explore the real world with a ground robot</title>
		<link>https://robohub.org/recon-learning-to-explore-the-real-world-with-a-ground-robot/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Tue, 09 Nov 2021 09:53:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<category><![CDATA[research]]></category>
		<guid isPermaLink="false">http://bair.berkeley.edu/blog/2021/11/03/recon/</guid>

					<description><![CDATA[










    
    
    An example of our method deployed on a Clearpath Jackal ground robot (left) exploring a suburban environment to find a visual target (inset). (Right) Egocentric observations of the robot.
    


Imagine you’re in an unfa...]]></description>
										<content:encoded><![CDATA[<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/recon/recon_blog_teaser.gif" alt="RECON Exploration Teaser" width="100%" /><i>An example of our method deployed on a Clearpath Jackal ground robot (left) exploring a suburban environment to find a visual target (inset). (Right) Egocentric observations of the robot.<br />
    </i>
</p>
<p>Imagine you’re in an unfamiliar neighborhood with no house numbers and I give you a photo that I took a few days ago of my house, which is not too far away. If you tried to find my house, you might follow the streets and go around the block looking for it. You might take a few wrong turns at first, but eventually you would locate my house. In the process, you would end up with a mental map of my neighborhood. The next time you’re visiting, you will likely be able to navigate to my house right away, without taking any wrong turns.</p>
<p>Such exploration and navigation behavior is easy for humans. What would it take for a robotic learning algorithm to enable this kind of intuitive navigation capability? To build a robot capable of exploring and navigating like this, we need to learn from diverse prior datasets in the real world. While it’s possible to collect a large amount of data from demonstrations, or even with randomized exploration, learning meaningful exploration and navigation behavior from this data can be challenging – the robot needs to generalize to unseen neighborhoods, recognize visual and dynamical similarities across scenes, and learn a representation of visual observations that is robust to distractors like weather conditions and obstacles. Since such factors can be hard to model and transfer from simulated environments, we tackle these problems by teaching the robot to explore using only real-world data.</p>
<p><span id="more-202383"></span></p>
<p>Formally, we studied the problem of <em>goal-directed</em> exploration for <em>visual</em> navigation in <em>novel</em> environments. A robot is tasked with navigating to a goal location <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-30a79c32f18567063fe44716929e7ced_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#71;" title="Rendered by QuickLaTeX.com" height="12" width="14" style="vertical-align: 0px;"/>, specified by an image <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-94746f680243719d02495da9fa7ca6cb_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#111;&#95;&#71;" title="Rendered by QuickLaTeX.com" height="11" width="20" style="vertical-align: -3px;"/> taken at <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-30a79c32f18567063fe44716929e7ced_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#71;" title="Rendered by QuickLaTeX.com" height="12" width="14" style="vertical-align: 0px;"/>. Our method uses an offline dataset of trajectories, over 40 hours of interactions in the real-world, to learn navigational affordances and builds a compressed representation of perceptual inputs. We deploy our method on a mobile robotic system in industrial and recreational outdoor areas around the city of Berkeley. RECON can discover a new goal in a previously unexplored environment in under 10 minutes, and in the process build a “mental map” of that environment that allows it to then reach goals again in just 20 seconds. Additionally, we make this real-world offline dataset publicly available for use in future research.</p>
<h2 id="rapid-exploration-controllers-for-outcome-driven-navigation">Rapid Exploration Controllers for Outcome-driven Navigation</h2>
<p>RECON, or <strong>R</strong>apid <strong>E</strong>xploration <strong>C</strong>ontrollers for <strong>O</strong>utcome-driven <strong>N</strong>avigation, explores new environments by “imagining” potential goal images and attempting to reach them. This exploration allows RECON to incrementally gather information about the new environment.</p>
<p>Our method consists of two components that enable it to explore new environments. The first component is a learned representation of goals. This representation ignores task-irrelevant distractors, allowing the agent to quickly adapt to novel settings. The second component is a topological graph. Our method learns both components using datasets or real-world robot interactions gathered in prior work. Leveraging such large datasets allows our method to generalize to new environments and scale beyond the original dataset.</p>
<h3 id="learning-to-represent-goals">Learning to Represent Goals</h3>
<p></p>
<p>A useful strategy to learn complex goal-reaching behavior in an unsupervised manner is for an agent to set its own goals, based on its capabilities, and attempt to reach them. <a href="https://pubmed.ncbi.nlm.nih.gov/15811218/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">In fact</a>, humans are very proficient at setting abstract goals for themselves in an effort to learn diverse skills. <a href="https://arxiv.org/abs/1807.04742" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Recent progress</a> in reinforcement learning and robotics has also shown that teaching agents to set its own goals by “imagining” them can result in learning of impressive unsupervised goal-reaching skills. To be able to “imagine”, or sample, such goals, we need to build a prior distribution over the goals seen during training.</p>
<p>For our case, where goals are represented by high-dimensional images, how should we sample goals for exploration? Instead of explicitly sampling goal images, we instead have the agent learn a compact representation of latent goals, allowing us to perform exploration by sampling new latent goal <em>representations</em>, rather than by sampling images. This representation of goals is learned from context-goal pairs previously seen by the robot. We use a <a href="https://arxiv.org/abs/1612.00410" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">variational information bottleneck</a> to learn these representations because it provides two important properties. First, it learns representations that throw away irrelevant information, such as lighting and pixel noise. Second, the variational information bottleneck packs the representations together so that they look like a chosen prior distribution. This is useful because we can then sample imaginary representations by sampling from this prior distribution.</p>
<p>The architecture for learning a prior distribution for these representations is shown below. As the encoder and decoder are conditioned on the context, the representation <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-503d34d286423ce2aefd62bcb3be1157_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#90;&#95;&#116;&#94;&#103;" title="Rendered by QuickLaTeX.com" height="20" width="19" style="vertical-align: -5px;"/> only encodes information about <em>relative</em> location of the goal from the context – this allows the model to represent feasible goals. If, instead, we had a typical VAE (in which the input images are autoencoded), the samples from the prior over these representations would not necessarily represent goals that are reachable from the current state. This distinction is crucial when exploring new environments, where most states from the training environments are not valid goals.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/recon/recon_blog_architecture.png" alt="Architecture with a latent goal model" width="100%" /><i>The architecture for learning a prior over goals in RECON. The context-conditioned embedding learns to represent feasible goals.<br />
    </i>
</p>
<p>To understand the importance of learning this representation, we run a simple experiment where the robot is asked to explore in an undirected manner starting from the yellow circle in the figure below. We find that sampling representations from the learned prior greatly accelerates the diversity of exploration trajectories and allows a wider area to be explored. In the absence of a prior over previously seen goals, using random actions to explore the environment can be quite inefficient. Sampling from the prior distribution and attempting to reach these “imagined” goals allows RECON to explore the environment efficiently.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/recon/recon_blog_sampling.png" alt="Goal sampling with RECON" width="100%" /><i>Sampling from a learned prior allows the robot to explore 5 times faster than using random actions.<br />
    </i>
</p>
<h3 id="goal-directed-exploration-with-a-topological-memory">Goal-Directed Exploration with a Topological Memory</h3>
<p></p>
<p>We combine this goal sampling scheme with a topological memory to incrementally build a “mental map” of the new environment. This map provides an estimate of the exploration <em>frontier</em> as well as guidance for subsequent exploration. In a new environment, RECON encourages the robot to explore at the frontier of the map – while the robot is not at the frontier, RECON directs it to navigate to a previously seen subgoal at the frontier of the map.</p>
<p>At the frontier, RECON uses the learned goal representation to learn a prior over goals it can reliably navigate to and are thus, <em>feasible</em> to reach. RECON uses this goal representation to sample, or “imagine”, a feasible goal that helps it explore the environment. This effectively means that, when placed in a new environment, if RECON does not know where the target is, it “imagines” a suitable subgoal that it can drive towards to explore and collects information, until it believes it can reach the target goal image. This allows RECON to “search” for the goal in an unknown environment, all the while building up its mental map. Note that the objective of the topological graph is to build a compact map of the environment and encourage the robot to reach the frontier; it does not inform goal sampling once the robot is at the frontier.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/recon/recon_blog_exploration.png" alt="Illustration of the exploration algorithm" width="100%" /><i>Illustration of the exploration algorithm of RECON.<br />
    </i>
</p>
<h3 id="learning-from-diverse-real-world-data">Learning from Diverse Real-world Data</h3>
<p></p>
<p>We train these models in RECON entirely using offline data collected in a diverse range of outdoor environments. Interestingly, we were able to train this model using data collected for two independent projects in the <a href="https://sites.google.com/view/badgr" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">fall of 2019</a> and <a href="https://sites.google.com/view/ving-robot/home" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">spring of 2020</a>, and were successful in deploying the model to explore novel environments and navigate to goals during late 2020 and the spring of 2021. This offline dataset of trajectories consists of over 40 hours of data, including off-road navigation, driving through parks in Berkeley and Oakland, parking lots, sidewalks and more, and is an excellent example of noisy real-world data with visual distractors like lighting, seasons (rain, twilight etc.), dynamic obstacles etc. The dataset consists of a mixture of teleoperated trajectories (2-3 hours) and open-loop safety controllers programmed to collect random data in a self-supervised manner. This dataset presents an exciting benchmark for robotic learning in real-world environments due to the challenges posed by offline learning of control, representation learning from high-dimensional visual observations, generalization to out-of-distribution environments and test-time adaptation.</p>
<p>We are releasing this dataset publicly to support future research in machine learning from real-world interaction datasets, check out the <a href="https://sites.google.com/view/recon-robot/dataset" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">dataset page</a> for more information.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/recon/recon_blog_envs.png" alt="Sample environments from the offline dataset of trajectories" width="100%" /><i>We train from diverse offline data (top) and test in new environments (bottom).<br />
    </i>
</p>
<h3 id="recon-in-action">RECON in Action</h3>
<p></p>
<p>Putting these components together, let’s see how RECON performs when deployed in a park near Berkeley. Note that the robot has never seen images from this park before. We placed the robot in a corner of the park and provided a target image of a white cabin door. In the animation below, we see RECON exploring and successfully finding the desired goal. “Run 1” corresponds to the exploration process in a novel environment, guided by a user-specified target image on the left. After it finds the goal, RECON uses the mental map to distill its experience in the environment to find the shortest path for subsequent traversals. In “Run 2”, RECON follows this path to navigate directly to the goal without looking around.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/recon/recon_blog_overview.gif" alt="Animation showing RECON deployed in a novel environment" width="100%" /><i>In “Run 1”, RECON explores a new environment and builds a topological mental map. In “Run 2”, it uses this mental map to quickly navigate to a user-specified goal in the environment.<br />
    </i>
</p>
<p>An illustration of this two-step process from an overhead view is show below, showing the paths taken by the robot in subsequent traversals of the environment:</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/recon/recon_blog_overview_overhead.png" alt="Overhead view of the exploration experiment above" width="100%" /><i>(Left) The goal specified by the user. (Right) The path taken by the robot when exploring for the first time (shown in cyan) to build a mental map with nodes (shown in white), and the path it takes when revisiting the same goal using the mental map (shown in red).<br />
    </i>
</p>
<h2 id="deploying-in-novel-environments">Deploying in Novel Environments</h2>
<p>To evaluate the performance of RECON in novel environments, study its behavior under a range of perturbations and understand the contributions of its components, we run extensive real-world experiments in the hills of Berkeley and Richmond, which have a diverse terrain and a wide variety of testing environments.</p>
<p>We compare RECON to five baselines – <a href="https://arxiv.org/abs/1810.12894" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RND</a>, <a href="https://arxiv.org/abs/1901.10902" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">InfoBot</a>, <a href="https://arxiv.org/abs/1901.10902" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Active Neural SLAM</a>, <a href="https://arxiv.org/abs/1901.10902" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ViNG</a> and <a href="https://arxiv.org/abs/1810.02274" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Episodic Curiosity</a> – each trained on the same offline trajectory dataset as our method, and fine-tuned in the target environment with online interaction. Note that this data is collected from past environments and contains no data from the target environment. The figure below shows the trajectories taken by the different methods for one such environment.</p>
<p>We find that only RECON (and a variant) is able to successfully discover the goal in over 30 minutes of exploration, while all other baselines result in collision (see figure for an overhead visualization). We visualize successful trajectories discovered by RECON in four other environments below.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/recon/recon_blog_overhead.png" alt="Overhead view comparing the different baselines in a novel environment" width="100%" /><br />
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/recon/recon_blog_trajectories.png" alt="Successful trajectories discovered by RECON in 4 different environments" width="100%" /><i>(Top) When comparing to other baselines, only RECON is able to successfully find the goal. (Bottom) Trajectories to goals in four other environments discovered by RECON.<br />
    </i>
</p>
<p>Quantitatively, we observe that our method finds goals over 50% faster than the best prior method; after discovering the goal and building a topological map of the environment, it can navigate to goals in that environment over 25% faster than the best alternative method.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/recon/recon_blog_exploration_barplot.png" alt="Quantitative results in novel environments" width="100%" /><i>Quantitative results in novel environments. RECON outperforms all baselines by over 50%.<br />
    </i>
</p>
<h3 id="exploring-non-stationary-environments">Exploring Non-Stationary Environments</h3>
<p></p>
<p>One of the important challenges in designing real-world robotic navigation systems is handling differences between training scenarios and testing scenarios. Typically, systems are developed in well-controlled environments, but are deployed in less structured environments. Further, the environments where robots are deployed often change over time, so tuning a system to perform well on a cloudy day might degrade performance on a sunny day. RECON uses explicit representation learning in attempts to handle this sort of non-stationary dynamics.</p>
<p>Our final experiment tested how changes in the environment affected the performance of RECON. We first had RECON explore a new “junkyard” to learn to reach a blue dumpster. Then, without any more supervision or exploration, we evaluated the learned policy when presented with <em>previously unseen obstacles</em> (trash cans, traffic cones, a car) and <em>weather conditions</em> (sunny, overcast, twilight). As shown below, RECON is able to successfully navigate to the goal in these scenarios, showing that the learned representations are invariant to visual distractors that do not affect the robot’s decisions to reach the goal.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/recon/recon_blog_obstacles.gif" alt="Robustness of RECON to novel obstacles" width="100%" /><br />
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/recon/recon_blog_weather.gif" alt="Robustness of RECON to variability in weather conditions" width="100%" /><i>First-person videos of RECON successfully navigating to a “blue dumpster” in the presence of novel obstacles (above) and varying weather conditions (below).<br />
    </i>
</p>
<h2 id="whats-next">What’s Next?</h2>
<p>The problem setup studied in this paper – using past experience to accelerate learning in a new environment – is reflective of several real-world robotics scenarios. RECON provides a robust way to solve this problem by using a combination of goal sampling and topological memory.</p>
<p>A mobile robot capable of reliably exploring and visually observing real-world environments can be a great tool for a wide variety of useful applications such as search and rescue, inspecting large offices or warehouses, finding leaks in oil pipelines or making rounds at a hospital, delivering mail in suburban communities. We demonstrated simplified versions of such applications <a href="https://sites.google.com/view/ving-robot/home" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">in an earlier project</a>, where the robot has prior experience in the deployment environment; RECON enables these results to scale beyond the training set of environments and results in a truly open-world learning system that can adapt to novel environments on deployment.</p>
<p>We are also releasing the aforementioned offline trajectory dataset, with hours of real-world interaction of a mobile ground robot in a variety of outdoor environments. We hope that this dataset can support future research in machine learning using real-world data for visual navigation applications. The dataset is also a rich source of sequential data from a multitude of sensors and can be used to test sequence prediction models including, but not limited to, video prediction, LiDAR, GPS etc. More information about the dataset can be found in the full-text article.</p>
<hr />
<p><em>This blog post is based on our paper <a href="https://arxiv.org/abs/2104.05859" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Rapid Exploration for Open-World Navigation with Latent Goal Models</a>, which will be presented as an Oral Talk at the 5th Annual Conference on Robot Learning in London, UK on November 8-11, 2021. You can find more information about our results and the dataset release on <a href="https://sites.google.com/view/recon-robot" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">the project page</a>.</em></p>
<p><em>Big thanks to Sergey Levine and Benjamin Eysenbach for helpful comments on an earlier draft of this article.</em></p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Making RL tractable by learning more informative reward functions: example-based control, meta-learning, and normalized maximum likelihood</title>
		<link>https://robohub.org/making-rl-tractable-by-learning-more-informative-reward-functions-example-based-control-meta-learning-and-normalized-maximum-likelihood/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sat, 30 Oct 2021 15:37:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<category><![CDATA[research]]></category>
		<guid isPermaLink="false">https://bair.berkeley.edu/blog/2021/10/22/mural/</guid>

					<description><![CDATA[














 Diagram of MURAL, our method for learning uncertainty-aware rewards for RL. After the user provides a few examples of desired outcomes, MURAL automatically infers a reward function that takes into account these examples and the age...]]></description>
										<content:encoded><![CDATA[<p style="text-align: left;"><img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_1.png" width="100%" /><br />
<i> Diagram of MURAL, our method for learning uncertainty-aware rewards for RL. After the user provides a few examples of desired outcomes, MURAL automatically infers a reward function that takes into account these examples and the agent’s uncertainty for each state.<br />
</i></p>
<p>Although reinforcement learning has shown success in domains <a href="https://arxiv.org/abs/1504.00702" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">such</a> <a href="https://arxiv.org/abs/2104.11203" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">as</a> <a href="https://arxiv.org/abs/1909.11652" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robotics</a>, chip <a href="https://arxiv.org/abs/2004.10746" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">placement</a> and <a href="https://www.nature.com/articles/s41586-019-1724-z" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">playing</a> <a href="https://arxiv.org/abs/1912.06680" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">video</a> <a href="https://www.nature.com/articles/nature16961" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">games</a>, it is usually intractable in its most general form. In particular, deciding when and how to visit new states in the hopes of learning more about the environment can be challenging, especially when the reward signal is uninformative. These questions of reward specification and exploration are closely connected — the more directed and “well shaped” a reward function is, the easier the problem of exploration becomes. The answer to the question of how to explore most effectively is likely to be closely informed by the particular choice of how we specify rewards.</p>
<p>For unstructured problem settings such as robotic manipulation and navigation — areas where RL holds substantial promise for enabling better real-world intelligent agents — reward specification is often the key factor preventing us from tackling more difficult tasks. The challenge of effective reward specification is two-fold: we require reward functions that can be specified in the real world without significantly instrumenting the environment, but also effectively guide the agent to solve difficult exploration problems. In our recent work, we address this challenge by designing a reward specification technique that naturally incentivizes exploration and enables agents to explore environments in a directed way.</p>
<h2 id="outcome-driven-rl-and-classifier-based-rewards">Outcome Driven RL and Classifier Based Rewards</h2>
<p>While RL in its most general form can be quite difficult to tackle, we can consider a more controlled set of subproblems which are more tractable while still encompassing a significant set of interesting problems. In particular, we consider a subclass of problems which has been referred to as <a href="https://proceedings.neurips.cc/paper/2018/file/c9319967c038f9b923068dabdf60cfe3-Paper.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">outcome driven RL</a>. In outcome driven RL problems, the agent is not simply tasked with exploring the environment until it chances upon reward, but instead is provided with examples of successful outcomes in the environment. These successful outcomes can then be used to infer a suitable reward function that can be optimized to solve the desired problems in new scenarios.</p>
<p>More concretely, in outcome driven RL problems, a human supervisor first provides a set of successful outcome examples <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-7d9e0c030375e9314c92cfa750db5876_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#123;&#115;&#95;&#103;&#94;&#105;&#125;&#95;&#123;&#105;&#61;&#49;&#125;&#94;&#78;" title="Rendered by QuickLaTeX.com" height="27" width="38" style="vertical-align: -8px;"/>, representing states in which the desired task has been accomplished. Given these outcome examples, a suitable reward function <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-fdd649aa12813f79e6432efeac936174_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#114;&#40;&#115;&#44;&#32;&#97;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="47" style="vertical-align: -5px;"/> can be inferred that encourages an agent to achieve the desired outcome examples. In many ways, this problem is analogous to that of inverse reinforcement learning, but only requires examples of successful states rather than full expert demonstrations.</p>
<p>When thinking about how to actually infer the desired reward function <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-fdd649aa12813f79e6432efeac936174_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#114;&#40;&#115;&#44;&#32;&#97;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="47" style="vertical-align: -5px;"/> from successful outcome examples <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-7d9e0c030375e9314c92cfa750db5876_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#123;&#115;&#95;&#103;&#94;&#105;&#125;&#95;&#123;&#105;&#61;&#49;&#125;&#94;&#78;" title="Rendered by QuickLaTeX.com" height="27" width="38" style="vertical-align: -8px;"/>, the simplest technique that comes to mind is to simply treat the reward inference problem as a classification problem &#8211; “Is the current state a successful outcome or not?” <a href="https://proceedings.neurips.cc/paper/2018/file/c9319967c038f9b923068dabdf60cfe3-Paper.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Prior</a> <a href="https://arxiv.org/abs/1904.07854" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">work</a> has implemented this intuition, inferring rewards by training a simple binary classifier to distinguish whether a particular state <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ae1901659f469e6be883797bfd30f4f8_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#115;" title="Rendered by QuickLaTeX.com" height="8" width="8" style="vertical-align: 0px;"/> is a successful outcome or not, using the set of provided goal states as positives, and all on-policy samples as negatives. The algorithm then assigns rewards to a particular state using the success probabilities from the classifier. This has been shown to have a close connection to the framework of inverse reinforcement learning.</p>
<p>Classifier-based methods provide a much more intuitive way to specify desired outcomes, removing the need for hand-designed reward functions or demonstrations:</p>
<p style="text-align: left;"><img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_2.png" width="100%" /></p>
<p>These classifier-based methods have achieved promising results on robotics tasks such as fabric placement, mug pushing, bead and screw manipulation, and more. However, these successes tend to be limited to simple shorter-horizon tasks, where relatively little exploration is required to find the goal.</p>
<h2 id="whats-missing">What’s Missing?</h2>
<p>Standard success classifiers in RL suffer from the key issue of overconfidence, which prevents them from providing useful shaping for hard exploration tasks. To understand why, let’s consider a toy 2D maze environment where the agent must navigate in a zigzag path from the top left to the bottom right corner. During training, classifier-based methods would label all on-policy states as negatives and user-provided outcome examples as positives. A typical neural network classifier would easily assign success probabilities of 0 to all visited states, resulting in uninformative rewards in the intermediate stages when the goal has not been reached.</p>
<p>Since such rewards would not be useful for guiding the agent in any particular direction, prior works tend to regularize their classifiers using methods like weight decay or mixup, which allow for more smoothly increasing rewards as we approach the successful outcome states. However, while this works on many shorter-horizon tasks, such methods can actually produce very misleading rewards. For example, on the 2D maze, a regularized classifier would assign relatively high rewards to states on the opposite side of the wall from the true goal, since they are close to the goal in x-y space. This causes the agent to get stuck in a local optima, never bothering to explore beyond the final wall!</p>
<p style="text-align: left;"><img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_3.png" width="100%" /></p>
<p>In fact, this is exactly what happens in practice:</p>
<p style="text-align: left;"><img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_4.gif" width="70%" /></p>
<h2 id="uncertainty-aware-rewards-through-cnml">Uncertainty-Aware Rewards through CNML</h2>
<p>As discussed above, the key issue with unregularized success classifiers for RL is overconfidence — by immediately assigning rewards of 0 to all visited states, we close off many paths that might eventually lead to the goal. Ideally, we would like our classifier to have an appropriate notion of uncertainty when outputting success probabilities, so that we can avoid excessively low rewards without suffering from the misleading local optima that result from regularization.</p>
<p><strong>Conditional Normalized Maximum Likelihood (CNML)</strong></p>
<p>One method particularly well-suited for this task is Conditional Normalized Maximum Likelihood (CNML). The concept of normalized maximum likelihood (NML) has typically been used in the Bayesian inference literature for model selection, to implement the minimum description length principle. In more recent work, NML has been adapted to the conditional setting to produce models that are much better calibrated and maintain a <a href="https://arxiv.org/abs/1812.09520" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">notion</a> of <a href="https://arxiv.org/abs/2011.02696" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">uncertainty</a>, while achieving optimal worst case classification regret. Given the challenges of overconfidence described above, this is an ideal choice for the problem of reward inference.</p>
<p>Rather than simply training models via maximum likelihood, CNML performs a more complex inference procedure to produce likelihoods for any point that is being queried for its label. Intuitively, CNML constructs a set of different maximum likelihood problems by labeling a particular query point <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ede05c264bba0eda080918aaa09c4658_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#120;" title="Rendered by QuickLaTeX.com" height="8" width="10" style="vertical-align: 0px;"/> with every possible label value that it might take, then outputs a final prediction based on how easily it was able to adapt to each of those proposed labels given the entire dataset observed thus far. Given a particular query point <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ede05c264bba0eda080918aaa09c4658_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#120;" title="Rendered by QuickLaTeX.com" height="8" width="10" style="vertical-align: 0px;"/>, and a prior dataset <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-4dcaa729b4b1f9b287b123cbf4415d4e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#68;&#125;&#32;&#61;&#32;&#92;&#108;&#101;&#102;&#116;&#91;&#120;&#95;&#48;&#44;&#32;&#121;&#95;&#48;&#44;&#32;&#8230;&#32;&#120;&#95;&#78;&#44;&#32;&#121;&#95;&#78;&#92;&#114;&#105;&#103;&#104;&#116;&#93;" title="Rendered by QuickLaTeX.com" height="18" width="172" style="vertical-align: -5px;"/>, CNML solves k different maximum likelihood problems and normalizes them to produce the desired label likelihood <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-93236510c2256f80b0a2fe8b81bca719_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#112;&#40;&#121;&#32;&#92;&#109;&#105;&#100;&#32;&#120;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="57" style="vertical-align: -5px;"/>, where <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3422b6bb5c160593658b7c39425d9880_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#107;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: 0px;"/> represents the number of possible values that the label may take. Formally, given a model <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-a7ee323bc5a3f73ad5e066b13bed5504_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#102;&#40;&#120;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="34" style="vertical-align: -5px;"/>, loss function <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-05327e4ac2b00755e99f445b74386a3e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;" title="Rendered by QuickLaTeX.com" height="12" width="12" style="vertical-align: 0px;"/>, training dataset <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-604919a910da61f809c93c64c0740a83_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#68;&#125;" title="Rendered by QuickLaTeX.com" height="12" width="14" style="vertical-align: 0px;"/> with classes <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-8f8851e49bba90cbe0336d6564950084_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#67;&#125;&#95;&#49;&#44;&#32;&#8230;&#44;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#67;&#125;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="16" width="72" style="vertical-align: -4px;"/>, and a new query point <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-60477a17b486f8b1f4fe1d1a4f40f1c5_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#120;&#95;&#113;" title="Rendered by QuickLaTeX.com" height="14" width="17" style="vertical-align: -6px;"/>, CNML solves the following <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3422b6bb5c160593658b7c39425d9880_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#107;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: 0px;"/> maximum likelihood problems:</p>
<p class="ql-center-displayed-equation" style="line-height: 26px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-f23601c051d52449da93909391be7caa_l3.png" height="26" width="272" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#92;&#116;&#104;&#101;&#116;&#97;&#95;&#105;&#32;&#61;&#32;&#92;&#116;&#101;&#120;&#116;&#123;&#97;&#114;&#103;&#125;&#92;&#109;&#97;&#120;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#69;&#125;&#95;&#123;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#68;&#125;&#32;&#92;&#99;&#117;&#112;&#32;&#40;&#120;&#95;&#113;&#44;&#32;&#67;&#95;&#105;&#41;&#125;&#92;&#108;&#101;&#102;&#116;&#91;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#40;&#102;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#125;&#40;&#120;&#41;&#44;&#32;&#121;&#41;&#92;&#114;&#105;&#103;&#104;&#116;&#93;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p>It then generates predictions for each of the <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3422b6bb5c160593658b7c39425d9880_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#107;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: 0px;"/> classes using their corresponding models, and normalizes the results for its final output:</p>
<p class="ql-center-displayed-equation" style="line-height: 70px;"><span class="ql-right-eqno"> &nbsp; </span><span class="ql-left-eqno"> &nbsp; </span><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ade6f974bec34a8a39b292fc099a2bee_l3.png" height="70" width="199" class="ql-img-displayed-equation quicklatex-auto-format" alt="&#92;&#91;&#112;&#95;&#92;&#116;&#101;&#120;&#116;&#123;&#67;&#78;&#77;&#76;&#125;&#40;&#67;&#95;&#105;&#124;&#120;&#41;&#32;&#61;&#32;&#92;&#102;&#114;&#97;&#99;&#123;&#102;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#95;&#105;&#125;&#40;&#120;&#41;&#125;&#123;&#92;&#115;&#117;&#109;&#32;&#92;&#108;&#105;&#109;&#105;&#116;&#115;&#95;&#123;&#106;&#61;&#49;&#125;&#94;&#107;&#32;&#102;&#95;&#123;&#92;&#116;&#104;&#101;&#116;&#97;&#95;&#106;&#125;&#40;&#120;&#41;&#125;&#92;&#93;" title="Rendered by QuickLaTeX.com"/></p>
<p style="text-align: left;"><img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_7.png" width="100%" /><br />
<i>Comparison of outputs from a standard classifier and a CNML classifier. CNML outputs more conservative predictions on points that are far from the training distribution, indicating uncertainty about those points’ true outputs. (Credit: Aurick Zhou, BAIR Blog)</i></p>
<p>Intuitively, if the query point is farther from the original training distribution represented by D, CNML will be able to more easily adapt to any arbitrary label in <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-8f8851e49bba90cbe0336d6564950084_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#67;&#125;&#95;&#49;&#44;&#32;&#8230;&#44;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#67;&#125;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="16" width="72" style="vertical-align: -4px;"/>, making the resulting predictions closer to uniform. In this way, CNML is able to produce better calibrated predictions, and maintain a clear notion of uncertainty based on which data point is being queried.</p>
<p style="text-align: left;"><img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_8.png" width="100%" /></p>
<p><strong>Leveraging CNML-based classifiers for Reward Inference</strong></p>
<p>Given the above background on CNML as a means to produce better calibrated classifiers, it becomes clear that this provides us a straightforward way to address the overconfidence problem with classifier based rewards in outcome driven RL. By replacing a standard maximum likelihood classifier with one trained using CNML, we are able to capture a notion of uncertainty and obtain directed exploration for outcome driven RL. In fact, in the discrete case, CNML corresponds to imposing a uniform prior on the output space — in an RL setting, this is equivalent to using a count-based exploration bonus as the reward function. This turns out to give us a very appropriate notion of uncertainty in the rewards, and solves many of the exploration challenges present in classifier based RL.</p>
<p>However, we don’t usually operate in the discrete case. In most cases, we use expressive function approximators and the resulting representations of different states in the world share similarities. When a CNML based classifier is learned in this scenario, with expressive function approximation, we see that it can provide more than just task agnostic exploration. In fact, it can provide a directed notion of reward shaping, which guides an agent towards the goal rather than simply encouraging it to expand the visited region naively. As visualized below, CNML encourages exploration by giving optimistic success probabilities in less-visited regions, while also providing better shaping towards the goal.</p>
<p style="text-align: left;"><img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_9.png" width="100%" /></p>
<p>As we will show in our experimental results, this intuition scales to higher dimensional problems and more complex state and action spaces, enabling CNML based rewards to solve significantly more challenging tasks than is possible with typical classifier based rewards.</p>
<p>However, on closer inspection of the CNML procedure, a major challenge becomes apparent. Each time a query is made to the CNML classifier, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3422b6bb5c160593658b7c39425d9880_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#107;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: 0px;"/> different maximum likelihood problems need to be solved to convergence, then normalized to produce the desired likelihood. As the size of the dataset increases, as it naturally does in reinforcement learning, this becomes a prohibitively slow process. In fact, as seen in Table 1, RL with standard CNML based rewards takes around 4 hours to train a single epoch (1000 timesteps). Following this procedure blindly would take over a month to train a single RL agent, necessitating a more time efficient solution. This is where we find meta-learning to be a crucial tool.</p>
<h2 id="meta-learning-cnml-classifiers">Meta-Learning CNML Classifiers</h2>
<p>Meta-learning is a tool that has seen a lot of use cases in few-shot learning for image classification, learning quicker optimizers and even learning more efficient RL algorithms. In essence, the idea behind meta-learning is to leverage a set of “meta-training” tasks to learn a model (and often an adaptation procedure) that can very quickly adapt to a new task drawn from the same distribution of problems.</p>
<p>Meta-learning techniques are particularly well suited to our class of computational problems since it involves quickly solving multiple different maximum likelihood problems to evaluate the CNML likelihood. Each the maximum likelihood problems share significant similarities with each other, enabling a meta-learning algorithm to very quickly adapt to produce solutions for each individual problem. In doing so, meta-learning provides us an effective tool for producing estimates of normalized maximum likelihood significantly more quickly than possible before.</p>
<p style="text-align: left;"><img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_10.gif" width="100%" /></p>
<p>The intuition behind how to apply meta-learning to the CNML (meta-NML) can be understood by the graphic above. For a data-set of <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-5793832f979c2268e3694c246d53b1bb_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#78;" title="Rendered by QuickLaTeX.com" height="12" width="16" style="vertical-align: 0px;"/> points, meta-NML would first construct <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-13b020be327c3fc955f64b2d96c329b1_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#50;&#78;" title="Rendered by QuickLaTeX.com" height="12" width="25" style="vertical-align: 0px;"/> tasks, corresponding to the positive and negative maximum likelihood problems for each datapoint in the dataset. Given these constructed tasks as a (meta) training set, a <a href="https://arxiv.org/abs/1703.03400" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">meta</a>&#8211;<a href="https://arxiv.org/abs/1703.05175" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">learning</a> algorithm can be applied to learn a model that can very quickly be adapted to produce solutions to any of these <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-13b020be327c3fc955f64b2d96c329b1_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#50;&#78;" title="Rendered by QuickLaTeX.com" height="12" width="25" style="vertical-align: 0px;"/> maximum likelihood problems. Equipped with this scheme to very quickly solve maximum likelihood problems, producing CNML predictions around <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-94eb615ea4179120cbb0e2cf72077709_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#52;&#48;&#48;" title="Rendered by QuickLaTeX.com" height="12" width="27" style="vertical-align: 0px;"/>x faster than possible before. Prior work studied this problem from a Bayesian approach, but we found that it often scales poorly for the problems we considered.</p>
<p>Equipped with a tool for efficiently producing predictions from the CNML distribution, we can now return to the goal of solving outcome-driven RL with uncertainty aware classifiers, resulting in an algorithm we call MURAL.</p>
<h2 id="mural-meta-learning-uncertainty-aware-rewards-for-automated-reinforcement-learning">MURAL: Meta-Learning Uncertainty-Aware Rewards for Automated Reinforcement Learning</h2>
<p>To more effectively solve outcome driven RL problems, we incorporate meta-NML into the standard classifier based procedure as follows: After each epoch of RL, we sample a batch of <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-b170995d512c659d8668b4e42e1fef6b_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#110;" title="Rendered by QuickLaTeX.com" height="8" width="11" style="vertical-align: 0px;"/> points from the replay buffer and use them to construct <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-7ca9d3b4ff5edce754df0706e9af8e9c_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#50;&#110;" title="Rendered by QuickLaTeX.com" height="12" width="20" style="vertical-align: 0px;"/> meta-tasks. We then run <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-4868771cbc422b5818f85500909ce433_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#49;" title="Rendered by QuickLaTeX.com" height="12" width="7" style="vertical-align: 0px;"/> iteration of meta-training on our model.<br />
We assign rewards using NML, where the NML outputs are approximated using only one gradient step for each input point.</p>
<p>The resulting algorithm, which we call MURAL, replaces the classifier portion of standard classifier-based RL algorithms with a meta-NML model instead. Although meta-NML can only evaluate input points one at a time instead of in batches, it is substantially faster than naive CNML, and MURAL is still comparable in runtime to standard classifier-based RL, as shown in Table 1 below.</p>
<p style="text-align: left;"><img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_11.png" width="60%" /><br />
<i>Table 1. Runtimes for a single epoch of RL on the 2D maze task.</i></p>
<p>We evaluate MURAL on a variety of navigation and robotic manipulation tasks, which present several challenges including local optima and difficult exploration. MURAL solves all of these tasks successfully, outperforming prior classifier-based methods as well as standard RL with exploration bonuses.</p>
<p style="text-align: left;"><img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_12.gif" width="50%" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_13.gif" width="50%" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_14.gif" width="50%" /></p>
<p><img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_15.gif" width="50%" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_16.gif" width="50%" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_17.gif" width="50%" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_18.gif" width="50%" /><br />
<i>Visualization of behaviors learned by MURAL. MURAL is able to perform a variety of behaviors in navigation and manipulation tasks, inferring rewards from outcome examples.</i></p>
<p style="text-align: left;"><img decoding="async" src="https://bair.berkeley.edu/static/blog/mural/MURAL_19.png" width="100%" /><br />
<i>Quantitative comparison of MURAL to baselines. MURAL is able to outperform baselines which perform task-agnostic exploration, standard maximum likelihood classifiers.</i></p>
<p>This suggests that using meta-NML based classifiers for outcome driven RL provides us an effective way to provide rewards for RL problems, providing benefits both in terms of exploration and directed reward shaping.</p>
<h2 id="takeaways">Takeaways</h2>
<p>In conclusion, we showed how outcome driven RL can define a class of more tractable RL problems. Standard methods using classifiers can often fall short in these settings as they are unable to provide any benefits of exploration or guidance towards the goal. Leveraging a scheme for training uncertainty aware classifiers via conditional normalized maximum likelihood allows us to more effectively solve this problem, providing benefits in terms of exploration and reward shaping towards successful outcomes. The general principles defined in this work suggest that considering tractable approximations to the general RL problem may allow us to simplify the challenge of reward specification and exploration in RL while still encompassing a rich class of control problems.</p>
<hr />
<p><i> This post is based on the paper “<a href="https://arxiv.org/abs/2107.07184" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">MURAL: Meta-Learning Uncertainty-Aware Rewards for Outcome-Driven Reinforcement Learning</a>”, which was presented at ICML 2021. You can see results <a href="https://sites.google.com/view/mural-rl" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">on our website</a>, and we <a href="https://github.com/mural-rl/mural" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">provide code</a> to reproduce our experiments.</i></p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>What can I do here? Learning new skills by imagining visual affordances</title>
		<link>https://robohub.org/what-can-i-do-here-learning-new-skills-by-imagining-visual-affordances/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Mon, 27 Sep 2021 12:45:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<category><![CDATA[research]]></category>
		<guid isPermaLink="false">https://robohub.org/what-can-i-do-here-learning-new-skills-by-imagining-visual-affordances/</guid>

					<description><![CDATA[<!--
These are comments in HTML. The above header text is needed to format the
title, authors, etc. The "example_post" is an example representative image (not
GIF) that we use for each post for tweeting (see below as well) and for the
emails to subscribers. Please provide this image (and any other images and
GIFs) in the blog to the BAIR Blog editors directly.

The text directly below gets tweets to work. Please adjust according to your
post.

The `static/blog` directory is a location on the blog server which permanently
stores the images/GIFs in BAIR Blog posts. Each post has a subdirectory under
this for its images (titled `example_post` here, please change).

Keeping the post visbility as False will mean the post is only accessible if
you know the exact URL.

You can also turn on Disqus comments, but we recommend disabling this feature.
-->

<!-- twitter -->
<!--
The actual text for the post content appears below.  Text will appear on the
homepage, i.e., https://bair.berkeley.edu/blog/ but we only show part of the
posts on the homepage. The rest is accessed via clicking 'Continue'. This is
enforced with the `more` excerpt separator.
--><p>How do humans become so skillful? Well, initially we are not, but from infancy, we discover and practice increasingly complex skills through self-supervised play. But this play is not random - the child development literature suggests that infants use their prior experience to conduct directed exploration of affordances like movability, suckability, graspability, and digestibility through interaction and sensory feedback. This type of affordance directed exploration allows infants to learn both what can be done in a given environment and how to do it. Can we instantiate an analogous strategy in a robotic learning system?</p>

<video autoplay="" loop="" muted="" width="80%"></video><p>On the left we see videos from a prior dataset collected with a robot accomplishing various tasks such as drawer opening and closing, as well as grasping and relocating objects. On the right we have a lid that the robot has never seen before. The robot has been granted a short period of time to practice with the new object, after which it will be given a goal image and tasked with making the scene match this image. How can the robot rapidly learn to manipulate the environment and grasp this lid without any external supervision?</p>

<video autoplay="" loop="" muted="" width="80%"></video><!--more--><p>To do so, we face several challenges. When a robot is dropped in a new environment, it must be able to use its prior knowledge to think of potentially useful behaviors that the environment affords. Then, the robot has to be able to actually practice these behaviors informatively. To now improve itself in the new environment, the robot must then be able to evaluate its own success somehow without an externally provided reward.</p>

<p>If we can overcome these challenges reliably, we open the door for a powerful cycle in which our agents use prior experience to collect high quality interaction data, which then grows their prior experience even further, continuously enhancing their potential utility!</p>

<h1>VAL: Visuomotor Affordance Learning</h1>

<p>Our method, Visuomotor Affordance Learning, or VAL, addresses these challenges. In VAL, we begin by assuming access to a prior dataset of robots demonstrating affordances in various environments. From here, VAL enters an offline phase which uses this information to learn 1) a generative model for imagining useful affordances in new environments,  2) a strong offline policy for effective exploration of these affordances, and 3) a self-evaluation metric for improving this policy. Finally, VAL is ready for it&#8217;s online phase. The agent is dropped in a new environment and can now use these learned capabilities to conduct self-supervised finetuning. The whole framework is illustrated in the figure below. Next, we will go deeper into the technical details of the offline and online phase.</p>

<video autoplay="" loop="" muted="" width="60%"></video><h1>VAL: Offline Phase</h1>

<p>Given a prior dataset demonstrating the affordances of various environments, VAL digests this information in three offline steps: representation learning to handle high dimensional real world data, affordance learning to enable self-supervised practice in unknown environments, and behavior learning to attain a high performance initial policy which accelerates online learning efficiency.</p>

<p><img src="https://bair.berkeley.edu/static/blog/val/image4.gif" alt="alt text"></p>

<p><strong>1.</strong> First, VAL learns a low representation of this data using a Vector Quantized Variational Auto-encoder or VQVAE. This process reduces our 48x48x3 images into a 144 dimensional latent space.</p>

<p>Distances in this latent space are meaningful, paving the way for our crucial mechanism of self-evaluating success. Given the current image s and goal image g, we encode both into the latent space, and threshold their distance to obtain a reward.</p>

<p>Later on, we will also use this representation as the latent space for our policy and Q function.</p>

<video autoplay="" loop="" muted=""></video><p><strong>2.</strong> Next, VAL learn an affordance model by training a PixelCNN in the latent space to the learn the distribution of reachable states conditioned on an image from the environment. This is done by maximizing the likelihood of the data,
$p(s_n &#124; s_0)$. We use this affordance model for directed exploration and for relabeling goals.</p>

<p>The affordance model is illustrated in the figure right. On the bottom left of the figure, we see that the conditioning image contains a pot, and the decoded latent goals on the upper right show the lid in different locations. These coherent goals will allow the robot to perform coherent exploration.</p>

<p><img src="https://bair.berkeley.edu/static/blog/val/image6.jpg" alt="alt text"></p>

<p><strong>3.</strong> Last in the offline phase, VAL must learn behaviors from the offline data, which it can then improve upon later with extra online, interactive data collection.</p>

<p>To accomplish this, we train a goal conditioned policy on the prior dataset using Advantage Weighted Actor Critic, an algorithm specifically designed for training offline and being amenable to online fine-tuning.</p>

<h1>VAL: Online Phase</h1>

<video autoplay="" loop="" muted=""></video><p>Now, when VAL is placed in an unseen environment, it uses its prior knowledge to imagine visual representations of useful affordances, collects helpful interaction data by trying to achieve these affordances, updates its parameters using its self-evaluation metric, and repeats the process all over again.</p>

<p>In this real example, on the left we see the initial state of the environment, which affords opening the drawer as well as other tasks.</p>

<p>In step 1, the affordance model samples a latent goal. By decoding the goal (using the VQVAE decoder, which is never actually used during RL because we operate entirely in the latent space), we can see the affordance is to open a drawer.</p>

<p>In step 2, we roll out the trained policy with the sampled goal. We see it successfully opens the drawer, in fact going too far and pulling the drawer all the way out. But this provides extremely useful interaction for the RL algorithm to further fine-tune on and perfect its policy.</p>

<p>After online finetuning is complete, we can now evaluate the robot on its ability to achieve the corresponding unseen goal images for each environment.</p>

<h1>Real World Evaluation</h1>

<p><img src="https://bair.berkeley.edu/static/blog/val/image8.jpg" alt="alt text"></p>

<p>We evaluate our method in five real-world test environments, and assess VAL on its ability to achieve a specific task the environment affords before and after <strong>five minutes</strong> of unsupervised fine-tuning.</p>

<p>Each test environment consists of at least one unseen interaction object, and two randomly sampled distractor objects. For instance, while there is opening and closing drawers in the training data, the new drawers have unseen handles.</p>

<video autoplay="" loop="" muted="" width="80%"></video><p>In every case, we begin with the offline trained policy, which solves the task inconsistently. Then, we collect more experience using our affordance model to sample goals. Finally, we evaluate the fine-tuned policy, which consistently solves the task.</p>

<p>We find that in each of these environments, VAL consistently demonstrates effective zero-shot generalization after offline training, followed by rapid improvement with its affordance-directed fine-tuning scheme. Meanwhile, prior self-supervised methods barely improve upon poor zero-shot performance in these new environments. These exciting results illustrate the potential that approaches like VAL possess for enabling robots to successfully operate far beyond the limited factory setting in which they are used to now.</p>

<p>Our dataset of 2,500 high quality robot interaction trajectories, covering 20 drawer handles, 20 pot handles, 60 toys, and 60 distractor objects, <a href="https://sites.google.com/view/val-rl/datasets" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">is now publicly available on our website</a>.</p>

<h1>Simulated Evaluation and Code</h1>

<p>For further analysis, we run VAL in a procedurally generated, multi-task environment with visual and dynamic variation. Which objects are in the scene, their colors, and their positions are randomized per environment. The agent can use handles to open drawers, grasp objects to relocate them, press buttons to unlock compartments, and so on.</p>

<p>The robot is given a prior dataset spanning various environments, and is evaluated on its ability to fine-tune on the following test environments.</p>

<p>Again, given a single off-policy dataset, our method quickly learns advanced manipulation skills including grasping, drawer opening, re-positioning, and tool usage for a diverse set of novel objects.</p>

<p>The environments and algorithm code are available; please see our <a href="https://github.com/anair13/rlkit/tree/master/examples/val" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">code repository</a>.</p>

<video autoplay="" loop="" muted="" width="80%"></video><h1>Future Work</h1>

<p>Like deep learning in domains such as computer vision and natural language processing which have been driven by large datasets and generalization, robotics will likely require learning from a similar scale of data. Because of this, improvements in offline reinforcement learning will be critical for enabling robots to take advantage of large prior datasets. Furthermore, these offline policies will need either rapid non-autonomous finetuning or entirely autonomous finetuning for real world deployment to be feasible. Lastly, once robots are operating on their own, we will have access to a continuous stream of new data, stressing both the importance and value of lifelong learning algorithms.</p>

<hr><p><i>This post is based on the paper &#8220;What Can I Do Here? Learning New Skills by Imagining Visual Affordances&#8221;, which was presented at the International Conference on Robotics and Automation (ICRA), 2021. You
can see results <a href="https://sites.google.com/view/val-rl" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">on our website</a>, and we <a href="https://github.com/anair13/rlkit/tree/master/examples/val" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">provide code</a> to to reproduce
our experiments.</i></p>]]></description>
										<content:encoded><![CDATA[<img decoding="async" src="https://robohub.org/wp-content/uploads/2021/09/IMG_20210927_144303-1024x550.jpg" alt="" width="1024" height="550" class="alignnone size-large wp-image-201493" srcset="https://robohub.org/wp-content/uploads/2021/09/IMG_20210927_144303-1024x550.jpg 1024w, https://robohub.org/wp-content/uploads/2021/09/IMG_20210927_144303-425x228.jpg 425w, https://robohub.org/wp-content/uploads/2021/09/IMG_20210927_144303-768x413.jpg 768w, https://robohub.org/wp-content/uploads/2021/09/IMG_20210927_144303.jpg 1295w" sizes="(max-width: 1024px) 100vw, 1024px" />
<p>How do humans become so skillful? Well, initially we are not, but from infancy, we discover and practice increasingly complex skills through self-supervised play. But this play is not random &#8211; the child development literature suggests that infants use their prior experience to conduct directed exploration of affordances like movability, suckability, graspability, and digestibility through interaction and sensory feedback. This type of affordance directed exploration allows infants to learn both what can be done in a given environment and how to do it. Can we instantiate an analogous strategy in a robotic learning system? <span id="more-201460"></span></p>
<p><video autoplay="" loop="" muted="" playsinline="" width="80%" style="display:block; margin: 0 auto;"><source src="https://bair.berkeley.edu/static/blog/val/image1.webm" type="video/webm" /><source src="https://bair.berkeley.edu/static/blog/val/image1.mp4" type="video/mp4" /></video></p>
<p>On the left we see videos from a prior dataset collected with a robot accomplishing various tasks such as drawer opening and closing, as well as grasping and relocating objects. On the right we have a lid that the robot has never seen before. The robot has been granted a short period of time to practice with the new object, after which it will be given a goal image and tasked with making the scene match this image. How can the robot rapidly learn to manipulate the environment and grasp this lid without any external supervision?</p>
<p><video autoplay="" loop="" muted="" playsinline="" width="80%" style="display:block; margin: 0 auto;"><source src="https://bair.berkeley.edu/static/blog/val/image2.webm" type="video/webm" /><source src="https://bair.berkeley.edu/static/blog/val/image2.mp4" type="video/mp4" /></video></p>
<p><!--more--></p>
<p>To do so, we face several challenges. When a robot is dropped in a new environment, it must be able to use its prior knowledge to think of potentially useful behaviors that the environment affords. Then, the robot has to be able to actually practice these behaviors informatively. To now improve itself in the new environment, the robot must then be able to evaluate its own success somehow without an externally provided reward.</p>
<p>If we can overcome these challenges reliably, we open the door for a powerful cycle in which our agents use prior experience to collect high quality interaction data, which then grows their prior experience even further, continuously enhancing their potential utility!</p>
<h1 id="val-visuomotor-affordance-learning">VAL: Visuomotor Affordance Learning</h1>
<p>Our method, Visuomotor Affordance Learning, or VAL, addresses these challenges. In VAL, we begin by assuming access to a prior dataset of robots demonstrating affordances in various environments. From here, VAL enters an offline phase which uses this information to learn 1) a generative model for imagining useful affordances in new environments,  2) a strong offline policy for effective exploration of these affordances, and 3) a self-evaluation metric for improving this policy. Finally, VAL is ready for it’s online phase. The agent is dropped in a new environment and can now use these learned capabilities to conduct self-supervised finetuning. The whole framework is illustrated in the figure below. Next, we will go deeper into the technical details of the offline and online phase.</p>
<p><video autoplay="" loop="" muted="" playsinline="" width="60%" style="display:block; margin: 0 auto;"><source src="https://bair.berkeley.edu/static/blog/val/image3.webm" type="video/webm" /><source src="https://bair.berkeley.edu/static/blog/val/image3.mp4" type="video/mp4" /></video></p>
<style type="text/css">
.image-left {
  display: block;
  margin-left: auto;
  margin-right: auto;
  float: right;
  width: 40%;
  padding: 3%;
}
</style>
<h1 id="val-offline-phase">VAL: Offline Phase</h1>
<p>Given a prior dataset demonstrating the affordances of various environments, VAL digests this information in three offline steps: representation learning to handle high dimensional real world data, affordance learning to enable self-supervised practice in unknown environments, and behavior learning to attain a high performance initial policy which accelerates online learning efficiency.</p>
<img decoding="async" src="https://bair.berkeley.edu/static/blog/val/image4.gif" alt="alt text" class="image-left" />
<p><strong>1.</strong> First, VAL learns a low representation of this data using a Vector Quantized Variational Auto-encoder or VQVAE. This process reduces our 48x48x3 images into a 144 dimensional latent space.</p>
<p>Distances in this latent space are meaningful, paving the way for our crucial mechanism of self-evaluating success. Given the current image s and goal image g, we encode both into the latent space, and threshold their distance to obtain a reward.</p>
<p>Later on, we will also use this representation as the latent space for our policy and Q function.</p>
<p><video autoplay="" loop="" muted="" playsinline="" class="image-left"><source src="https://bair.berkeley.edu/static/blog/val/image5.webm" type="video/webm" /><source src="https://bair.berkeley.edu/static/blog/val/image5.mp4" type="video/mp4" /></video></p>
<p><strong>2.</strong> Next, VAL learn an affordance model by training a PixelCNN in the latent space to the learn the distribution of reachable states conditioned on an image from the environment. This is done by maximizing the likelihood of the data,<br />
$p(s_n | s_0)$. We use this affordance model for directed exploration and for relabeling goals.</p>
<p>The affordance model is illustrated in the figure right. On the bottom left of the figure, we see that the conditioning image contains a pot, and the decoded latent goals on the upper right show the lid in different locations. These coherent goals will allow the robot to perform coherent exploration.</p>
<img decoding="async" src="https://bair.berkeley.edu/static/blog/val/image6.jpg" alt="alt text" class="image-left" />
<p><strong>3.</strong> Last in the offline phase, VAL must learn behaviors from the offline data, which it can then improve upon later with extra online, interactive data collection.</p>
<p>To accomplish this, we train a goal conditioned policy on the prior dataset using Advantage Weighted Actor Critic, an algorithm specifically designed for training offline and being amenable to online fine-tuning.</p>
<h1 id="val-online-phase">VAL: Online Phase</h1>
<p><video autoplay="" loop="" muted="" playsinline="" class="image-left"><source src="https://bair.berkeley.edu/static/blog/val/image7.webm" type="video/webm" /><source src="https://bair.berkeley.edu/static/blog/val/image7.mp4" type="video/mp4" /></video></p>
<p>Now, when VAL is placed in an unseen environment, it uses its prior knowledge to imagine visual representations of useful affordances, collects helpful interaction data by trying to achieve these affordances, updates its parameters using its self-evaluation metric, and repeats the process all over again.</p>
<p>In this real example, on the left we see the initial state of the environment, which affords opening the drawer as well as other tasks.</p>
<p>In step 1, the affordance model samples a latent goal. By decoding the goal (using the VQVAE decoder, which is never actually used during RL because we operate entirely in the latent space), we can see the affordance is to open a drawer.</p>
<p>In step 2, we roll out the trained policy with the sampled goal. We see it successfully opens the drawer, in fact going too far and pulling the drawer all the way out. But this provides extremely useful interaction for the RL algorithm to further fine-tune on and perfect its policy.</p>
<p>After online finetuning is complete, we can now evaluate the robot on its ability to achieve the corresponding unseen goal images for each environment.</p>
<h1 id="real-world-evaluation">Real World Evaluation</h1>
<img decoding="async" src="https://bair.berkeley.edu/static/blog/val/image8.jpg" alt="alt text" class="image-left" />
<p>We evaluate our method in five real-world test environments, and assess VAL on its ability to achieve a specific task the environment affords before and after <strong>five minutes</strong> of unsupervised fine-tuning.</p>
<p>Each test environment consists of at least one unseen interaction object, and two randomly sampled distractor objects. For instance, while there is opening and closing drawers in the training data, the new drawers have unseen handles.</p>
<p><video autoplay="" loop="" muted="" playsinline="" width="80%" style="display:block; margin: 0 auto;"><source src="https://bair.berkeley.edu/static/blog/val/image9.webm" type="video/webm" /><source src="https://bair.berkeley.edu/static/blog/val/image9.mp4" type="video/mp4" /></video></p>
<p>In every case, we begin with the offline trained policy, which solves the task inconsistently. Then, we collect more experience using our affordance model to sample goals. Finally, we evaluate the fine-tuned policy, which consistently solves the task.</p>
<p>We find that in each of these environments, VAL consistently demonstrates effective zero-shot generalization after offline training, followed by rapid improvement with its affordance-directed fine-tuning scheme. Meanwhile, prior self-supervised methods barely improve upon poor zero-shot performance in these new environments. These exciting results illustrate the potential that approaches like VAL possess for enabling robots to successfully operate far beyond the limited factory setting in which they are used to now.</p>
<p>Our dataset of 2,500 high quality robot interaction trajectories, covering 20 drawer handles, 20 pot handles, 60 toys, and 60 distractor objects, <a href="https://sites.google.com/view/val-rl/datasets" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">is now publicly available on our website</a>.</p>
<h1 id="simulated-evaluation-and-code">Simulated Evaluation and Code</h1>
<p>For further analysis, we run VAL in a procedurally generated, multi-task environment with visual and dynamic variation. Which objects are in the scene, their colors, and their positions are randomized per environment. The agent can use handles to open drawers, grasp objects to relocate them, press buttons to unlock compartments, and so on.</p>
<p>The robot is given a prior dataset spanning various environments, and is evaluated on its ability to fine-tune on the following test environments.</p>
<p>Again, given a single off-policy dataset, our method quickly learns advanced manipulation skills including grasping, drawer opening, re-positioning, and tool usage for a diverse set of novel objects.</p>
<p>The environments and algorithm code are available; please see our <a href="https://github.com/anair13/rlkit/tree/master/examples/val" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">code repository</a>.</p>
<p><video autoplay="" loop="" muted="" playsinline="" width="80%" style="display:block; margin: 0 auto;"><source src="https://bair.berkeley.edu/static/blog/val/image10.webm" type="video/webm" /><source src="https://bair.berkeley.edu/static/blog/val/image10.mp4" type="video/mp4" /></video></p>
<h1 id="future-work">Future Work</h1>
<p>Like deep learning in domains such as computer vision and natural language processing which have been driven by large datasets and generalization, robotics will likely require learning from a similar scale of data. Because of this, improvements in offline reinforcement learning will be critical for enabling robots to take advantage of large prior datasets. Furthermore, these offline policies will need either rapid non-autonomous finetuning or entirely autonomous finetuning for real world deployment to be feasible. Lastly, once robots are operating on their own, we will have access to a continuous stream of new data, stressing both the importance and value of lifelong learning algorithms.</p>
<hr />
<p><i>This post is based on the paper “What Can I Do Here? Learning New Skills by Imagining Visual Affordances”, which was presented at the International Conference on Robotics and Automation (ICRA), 2021. You can see results <a href="https://sites.google.com/view/val-rl" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">on our website</a>, and we <a href="https://github.com/anair13/rlkit/tree/master/examples/val" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">provide code</a> to reproduce our experiments.</i></p>
]]></content:encoded>
					
		
		<enclosure url="https://bair.berkeley.edu/static/blog/val/image1.webm" length="173413" type="video/webm" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image1.mp4" length="195411" type="video/mp4" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image2.webm" length="1356238" type="video/webm" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image2.mp4" length="1298954" type="video/mp4" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image3.webm" length="183993" type="video/webm" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image3.mp4" length="165136" type="video/mp4" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image5.webm" length="26816" type="video/webm" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image5.mp4" length="25871" type="video/mp4" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image7.webm" length="304572" type="video/webm" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image7.mp4" length="335230" type="video/mp4" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image9.webm" length="1094458" type="video/webm" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image9.mp4" length="1359755" type="video/mp4" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image10.webm" length="73840" type="video/webm" />
<enclosure url="https://bair.berkeley.edu/static/blog/val/image10.mp4" length="76178" type="video/mp4" />

			</item>
		<item>
		<title>Maximum Entropy RL (Provably) Solves Some Robust RL Problems</title>
		<link>https://robohub.org/maximum-entropy-rl-provably-solves-some-robust-rl-problems/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Mon, 15 Mar 2021 10:00:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<category><![CDATA[manipulation]]></category>
		<category><![CDATA[research]]></category>
		<guid isPermaLink="false">https://robohub.org/maximum-entropy-rl-provably-solves-some-robust-rl-problems/</guid>

					<description><![CDATA[<!-- twitter -->
<p>Nearly all real-world applications of reinforcement learning involve some degree of shift between the training environment and the testing environment. However, prior work has observed that even small shifts in the environment cause most RL algorithms to perform <a href="https://arxiv.org/abs/1703.06907" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">markedly</a> <a href="https://arxiv.org/abs/1610.01283" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">worse</a>.
As we aim to scale reinforcement learning algorithms and apply them in the real world, it is increasingly important to learn policies that are robust to changes in the environment.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/robust_rl.gif" width="90%"><br><i><b>Robust reinforcement learning</b> maximizes reward on an adversarially-chosen environment.</i>
</p>

<p>Broadly, prior approaches to handling distribution shift in RL aim to maximize performance in either the average case or the worst case. The first set of approaches, such as domain randomization, train a policy on a distribution of environments, and optimize the average performance of the policy on these environments. While these methods have been successfully applied to a number of areas
(e.g., <a href="https://arxiv.org/abs/1804.09364" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">self-driving cars</a>, <a href="https://arxiv.org/abs/1804.10332" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robot locomotion</a> and <a href="https://arxiv.org/abs/1703.06907" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">manipulation</a>),
their success rests critically on the <a href="https://arxiv.org/abs/1910.07113" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">design of the distribution of environments</a>.
Moreover, policies that do well on average are not guaranteed to get high reward on every environment. The policy that gets the highest reward on average might get very low reward on a small fraction of environments. The second set of approaches, typically referred to as <strong>robust RL</strong>, focus on the worst-case scenarios. The aim is to find a policy that gets high reward on every environment within some set. Robust RL can equivalently be viewed as a <a href="https://www.youtube.com/watch?v=xfyK03MEZ9Q&#038;t=17093s" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">two-player game</a> between the policy and an environment adversary. The policy tries to get high reward, while the environment adversary tries to tweak the dynamics and reward function of the environment so that the policy gets lower reward. One important property of the robust approach is that, unlike domain randomization, it is invariant to the ratio of easy and hard tasks. Whereas robust RL always evaluates a policy on the most challenging tasks, domain randomization will predict that the policy is better if it is evaluated on a distribution of environments with more easy tasks.</p>

<!--more-->

<p>Prior work has suggested a number of algorithms for solving robust RL problems. Generally, these algorithms all follow the same recipe: take an existing RL algorithm and add some additional machinery on top to make it robust.
For example, <a href="https://www.ri.cmu.edu/pub_files/pub3/bagnell_james_2001_1/bagnell_james_2001_1.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robust value iteration</a> uses Q-learning as the base RL algorithm, and modifies the Bellman update by solving a convex optimization problem in the inner loop of each Bellman backup.
Similarly, <a href="http://proceedings.mlr.press/v70/pinto17a/pinto17a.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Pinto &#8216;17</a> uses TRPO as the base RL algorithm and periodically updates the environment based on the behavior of the current policy. These prior approaches are often difficult to implement and, even once implemented correctly, they requiring tuning of many additional hyperparameters. Might there be a simpler approach, an approach that does not require additional hyperparameters and additional lines of code to debug?</p>

<p>To answer this question, we are going to focus on a type of RL algorithm known as maximum entropy RL, or <strong>MaxEnt RL</strong> for short (<a href="https://proceedings.neurips.cc/paper/2006/file/d806ca13ca3449af72a1ea5aedbed26a-Paper.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Todorov &#8216;06</a>, <a href="http://www.roboticsproceedings.org/rss08/p45.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Rawlik &#8216;08</a>, <a href="https://www.cs.uic.edu/pub/Ziebart/Publications/thesis-bziebart.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Ziebart &#8216;10</a>).
MaxEnt RL is a slight variant of standard RL that aims to learn a policy that gets high reward while acting as randomly as possible; formally, MaxEnt maximizes the entropy of the policy. Some <a href="https://arxiv.org/abs/1812.11103" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">prior</a> <a href="https://openreview.net/forum?id=r1xPh2VtPB" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">work</a> has observed empirically that MaxEnt RL algorithms appear to be robust to some disturbances the environment.
To the best of our knowledge, no prior work has actually proven that MaxEnt RL is robust to environmental disturbances.</p>

<p>In a <a href="https://arxiv.org/abs/2103.06257" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recent paper</a>, we prove that every MaxEnt RL problem corresponds to maximizing a lower bound on a robust RL problem. Thus, when you run MaxEnt RL, you are implicitly solving a robust RL problem. Our analysis provides a theoretically-justified explanation for the empirical robustness of MaxEnt RL, and proves that <em>MaxEnt RL is itself a robust RL algorithm.</em>
In the rest of this post, we&#8217;ll provide some intuition into why MaxEnt RL should be robust and what sort of perturbations MaxEnt RL is robust to. We&#8217;ll also show some experiments demonstrating the robustness of MaxEnt RL.</p>

<h1>Intuition</h1>

<p>So, why would we expect MaxEnt RL to be robust to disturbances in the environment? Recall that MaxEnt RL trains policies to not only maximize reward, but to do so while acting as randomly as possible. In essence, the policy itself is injecting as much noise as possible into the environment, so it gets to &#8220;practice&#8221; recovering from disturbances. Thus, if the change in dynamics appears like just a disturbance in the original environment, our policy has already been trained on such data. Another way of viewing MaxEnt RL is as learning many different ways of solving the task (<a href="https://www.cs.uic.edu/pub/Ziebart/Publications/thesis-bziebart.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Kappen &#8216;05</a>). For example, let&#8217;s look at the task shown in videos below: we want the robot to push the white object to the green region. The top two videos show that standard RL always takes the shortest path to the goal, whereas MaxEnt RL takes many different paths to the goal. Now, let&#8217;s imagine that we add a new obstacle (red blocks) that wasn&#8217;t included during training. As shown in the videos in the bottom row, the policy learned by standard RL almost always collides with the obstacle, rarely reaching the goal. In contrast, the MaxEnt RL policy often chooses routes around the obstacle, continuing to reach the goal for a large fraction of trials.</p>

<p>
</p><table><tr><td>
  </td>
  <td>
    Standard RL
  </td>
  <td>
    MaxEnt RL
  </td>
</tr><tr><td>
    <p>Trained and evaluated without the obstacle:</p>
  </td>
  <td>
    <img src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/pusher_empty_standard.gif" width="100%"></td>
  <td>
    <img src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/pusher_empty_maxent.gif" width="100%"></td>
</tr><tr><td>
    <p>Trained without the obstacle, but evaluated with
    the obstacle:</p>
  </td>
  <td>
    <img src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/pusher_obstacle_standard.gif" width="100%"></td>
  <td>
    <img src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/pusher_obstacle_maxent.gif" width="100%"></td>
</tr></table><h1>Theory</h1>

<p>We now formally describe the technical results from the paper. The aim here is not to provide a full proof (see the paper Appendix for that), but instead to build some intuition for what the technical results say. Our main result is that, when you apply MaxEnt RL with some reward function and some dynamics, you are actually maximizing a lower bound on the robust RL objective. To explain this result, we must first define the MaxEnt RL objective:
$J_{MaxEnt}(\pi; p, r)$ is the entropy-regularized cumulative return of policy $\pi$ when evaluated using dynamics $p(s&#8217; \mid s, a)$ and reward function $r(s, a)$. While we will train the policy using one dynamics $p$, we will evaluate the policy on a different dynamics, $\tilde{p}(s&#8217; \mid s, a)$, chosen by the adversary. We can now formally state our main result as follows:</p>

<p>The left-hand-side is the robust RL objective. It says that the adversary gets
to choose whichever dynamics function $\tilde{p}(s&#8217; \mid s, a)$ makes our policy perform as poorly as
possible, subject to some constraints (as specified by the set $\tilde{\mathcal{P}}$).  On
the right-hand-side we have the MaxEnt RL objective (note that $\log T$ is a
constant, and the function $\exp(\cdots)$ is always increasing). Thus, this objective
says that a policy that has a high entropy-regularized reward (right hand-side)
is guaranteed to also get high reward when evaluated on an adversarially-chosen
dynamics.</p>

<p>The most important part of this equation is the set $\tilde{\mathcal{P}}$ of dynamics that
the adversary can choose from. Our analysis describes precisely how this set is
constructed and shows that, if we want a policy to be robust to a larger set of
disturbances, all we have to do is increase the weight on the entropy term and
decrease the weight on the reward term. Intuitively, the adversary must choose
dynamics that are &#8220;close&#8221; to the dynamics on which the policy was trained. For
example, in the special case where the dynamics are linear-Gaussian, this set
corresponds to all perturbations where the original expected next state and the
perturbed expected next state have a Euclidean distance less than $\epsilon$.</p>

<h1>More Experiments</h1>

<p>Our analysis predicts that MaxEnt RL should be robust to many types of
disturbances. The first set of videos in this post showed that MaxEnt RL is robust to 
static obstacles. MaxEnt RL is also robust to dynamic perturbations introduced in the
middle of an episode. To demonstrate this, we took the same robotic pushing task
and knocked the puck out of place in the middle of the episode. The videos below
show that the policy learned by MaxEnt RL is more robust at handling these
perturbations, as predicted by our analysis.</p>

<p>
</p><table><tr><td>
    <p>Standard RL</p>
  </td>
  <td>
    <p>MaxEnt RL</p>
  </td>
</tr><tr><td>
    <img src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/pusher_force_standard_v2_opt.gif" width="100%"></td>
  <td>
    <img src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/pusher_force_maxent_v2_opt.gif" width="100%"></td>
</tr></table><p><i>The policy learned by MaxEntRL is robust to dynamic perturbations of the puck (red frames).
</i></p>


<p>Our theoretical results suggest that, even if we optimize the environment
perturbations so the agent does as poorly as possible, MaxEnt RL policies will
still be robust. To demonstrate this capability, we trained both standard RL and
MaxEnt RL on a peg insertion task shown below. During evaluation, we changed the
position of the hole to try to make each policy fail. If we only moved the hole
position a little bit ($\le$ 1 cm), both policies always solved the task. However,
if we moved the hole position up to 2cm, the policy learned by standard RL
almost never succeeded in inserting the peg, while the MaxEnt RL policy
succeeded in 95% of trials. This experiment validates our
theoretical findings that MaxEnt really is robust to (bounded) adversarial
disturbances in the environment.</p>

<p>
</p><table><colgroup><col span="1"><col span="1"><col span="1"></colgroup><tr><td>
    <img src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/peg_standard_long.gif" width="100%"></td>
  <td>
    <img src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/peg_maxent_long.gif" width="100%"></td>
  <td>
    <img src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/peg_minimax_100.png" width="100%"></td>
</tr><tr><td>
    <p>Standard RL</p>
  </td>
  <td>
    <p>MaxEnt RL</p>
  </td>
  <td>
    <p>Evaluation on adversarial perturbations</p>
  </td>
</tr></table><p>
<i>MaxEnt RL is robust to adversarial perturbations of the hole (where the robot
inserts the peg).</i></p>


<h1>Conclusion</h1>

<p>In summary, <a href="https://arxiv.org/abs/2103.06257" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our paper</a> shows that a commonly-used type of RL algorithm, MaxEnt
RL, is already solving a robust RL problem. We do not claim that MaxEnt RL will
outperform purpose-designed robust RL algorithms. However, the striking
simplicity of MaxEnt RL compared with other robust RL algorithms suggests that
it may be an appealing alternative to practitioners hoping to equip their RL
policies with an ounce of robustness.</p>

<p><strong>Acknowledgements</strong>
Thanks to Gokul Swamy, Diba Ghosh, Colin Li, and Sergey Levine for feedback on drafts of this post,
and to Chloe Hsu and Daniel Seita for help with the blog.</p>

<hr><p>This post is based on the following paper:</p>

<ul><li><a href="https://arxiv.org/abs/2103.06257" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Maximum Entropy RL (Provably) Solves Some Robust RL Problems</a>. <br><a href="https://ben-eysenbach.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Benjamin Eysenbach</a> and <a href="https://people.eecs.berkeley.edu/~svlevine/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sergey Levine</a>.</li>
</ul>]]></description>
										<content:encoded><![CDATA[<p><strong>By Ben Eysenbach</strong></p>
<p>Nearly all real-world applications of reinforcement learning involve some degree of shift between the training environment and the testing environment. However, prior work has observed that even small shifts in the environment cause most RL algorithms to perform <a href="https://arxiv.org/abs/1703.06907" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">markedly</a> <a href="https://arxiv.org/abs/1610.01283" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">worse</a>. As we aim to scale reinforcement learning algorithms and apply them in the real world, it is increasingly important to learn policies that are robust to changes in the environment.</p>
<p style="text-align:center; float:right; width:50%; padding-left:15px;
padding-right:15px; margin-bottom:0px"><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/robust_rl.gif" width="90%" /><br />
<br />
<i><b>Robust reinforcement learning</b> maximizes reward on an adversarially-chosen environment.</i>
</p>
<p>Broadly, prior approaches to handling distribution shift in RL aim to maximize performance in either the average case or the worst case. The first set of approaches, such as domain randomization, train a policy on a distribution of environments, and optimize the average performance of the policy on these environments. While these methods have been successfully applied to a number of areas (e.g., <a href="https://arxiv.org/abs/1804.09364" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">self-driving cars</a>, <a href="https://arxiv.org/abs/1804.10332" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robot locomotion</a> and <a href="https://arxiv.org/abs/1703.06907" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">manipulation</a>), their success rests critically on the <a href="https://arxiv.org/abs/1910.07113" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">design of the distribution of environments</a>. Moreover, policies that do well on average are not guaranteed to get high reward on every environment. The policy that gets the highest reward on average might get very low reward on a small fraction of environments. The second set of approaches, typically referred to as <strong>robust RL</strong>, focus on the worst-case scenarios. The aim is to find a policy that gets high reward on every environment within some set. Robust RL can equivalently be viewed as a <a href="https://www.youtube.com/watch?v=xfyK03MEZ9Q&amp;t=17093s" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">two-player game</a> between the policy and an environment adversary. The policy tries to get high reward, while the environment adversary tries to tweak the dynamics and reward function of the environment so that the policy gets lower reward. One important property of the robust approach is that, unlike domain randomization, it is invariant to the ratio of easy and hard tasks. Whereas robust RL always evaluates a policy on the most challenging tasks, domain randomization will predict that the policy is better if it is evaluated on a distribution of environments with more easy tasks.</p>
<p>Prior work has suggested a number of algorithms for solving robust RL problems. Generally, these algorithms all follow the same recipe: take an existing RL algorithm and add some additional machinery on top to make it robust. For example, <a href="https://www.ri.cmu.edu/pub_files/pub3/bagnell_james_2001_1/bagnell_james_2001_1.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robust value iteration</a> uses Q-learning as the base RL algorithm, and modifies the Bellman update by solving a convex optimization problem in the inner loop of each Bellman backup. Similarly, <a href="http://proceedings.mlr.press/v70/pinto17a/pinto17a.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Pinto ‘17</a> uses TRPO as the base RL algorithm and periodically updates the environment based on the behavior of the current policy. These prior approaches are often difficult to implement and, even once implemented correctly, they requiring tuning of many additional hyperparameters. Might there be a simpler approach, an approach that does not require additional hyperparameters and additional lines of code to debug?</p>
<p>To answer this question, we are going to focus on a type of RL algorithm known as maximum entropy RL, or <strong>MaxEnt RL</strong> for short (<a href="https://proceedings.neurips.cc/paper/2006/file/d806ca13ca3449af72a1ea5aedbed26a-Paper.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Todorov ‘06</a>, <a href="http://www.roboticsproceedings.org/rss08/p45.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Rawlik ‘08</a>, <a href="https://www.cs.uic.edu/pub/Ziebart/Publications/thesis-bziebart.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Ziebart ‘10</a>). MaxEnt RL is a slight variant of standard RL that aims to learn a policy that gets high reward while acting as randomly as possible; formally, MaxEnt maximizes the entropy of the policy. Some <a href="https://arxiv.org/abs/1812.11103" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">prior</a> <a href="https://openreview.net/forum?id=r1xPh2VtPB" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">work</a> has observed empirically that MaxEnt RL algorithms appear to be robust to some disturbances the environment. To the best of our knowledge, no prior work has actually proven that MaxEnt RL is robust to environmental disturbances.</p>
<p>In a <a href="https://arxiv.org/abs/2103.06257" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recent paper</a>, we prove that every MaxEnt RL problem corresponds to maximizing a lower bound on a robust RL problem. Thus, when you run MaxEnt RL, you are implicitly solving a robust RL problem. Our analysis provides a theoretically-justified explanation for the empirical robustness of MaxEnt RL, and proves that <em>MaxEnt RL is itself a robust RL algorithm.</em> In the rest of this post, we’ll provide some intuition into why MaxEnt RL should be robust and what sort of perturbations MaxEnt RL is robust to. We’ll also show some experiments demonstrating the robustness of MaxEnt RL.</p>
<h1 id="intuition">Intuition</h1>
<p>So, why would we expect MaxEnt RL to be robust to disturbances in the environment? Recall that MaxEnt RL trains policies to not only maximize reward, but to do so while acting as randomly as possible. In essence, the policy itself is injecting as much noise as possible into the environment, so it gets to “practice” recovering from disturbances. Thus, if the change in dynamics appears like just a disturbance in the original environment, our policy has already been trained on such data. Another way of viewing MaxEnt RL is as learning many different ways of solving the task (<a href="https://www.cs.uic.edu/pub/Ziebart/Publications/thesis-bziebart.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Kappen ‘05</a>). For example, let’s look at the task shown in videos below: we want the robot to push the white object to the green region. The top two videos show that standard RL always takes the shortest path to the goal, whereas MaxEnt RL takes many different paths to the goal. Now, let’s imagine that we add a new obstacle (red blocks) that wasn’t included during training. As shown in the videos in the bottom row, the policy learned by standard RL almost always collides with the obstacle, rarely reaching the goal. In contrast, the MaxEnt RL policy often chooses routes around the obstacle, continuing to reach the goal for a large fraction of trials.</p>
<p style="text-align:center;">
<table>
<tr>
<td style="border-top:none; border-bottom:none">
  </td>
<td style="border-top:none; border-bottom:none; text-align:center; padding:0px">
    Standard RL
  </td>
<td style="border-top:none; border-bottom:none; text-align:center; padding:0px">
    MaxEnt RL
  </td>
</tr>
<tr>
<td style="border-top:none; border-bottom:none; padding:0px;
  vertical-align:middle;"></p>
<p style="text-align:center;">Trained and evaluated without the obstacle:</p>
</td>
<td style="border-top:none; border-bottom:none; padding:0px;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/pusher_empty_standard.gif" width="100%" />
  </td>
<td style="border-top:none; border-bottom:none; padding:0px;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/pusher_empty_maxent.gif" width="100%" />
  </td>
</tr>
<tr>
<td style="border-top:none; border-bottom:none; padding:0px;
  vertical-align:middle;"></p>
<p style="text-align:center;">Trained without the obstacle, but evaluated with<br />
    the obstacle:</p>
</td>
<td style="border-top:none; border-bottom:none; padding:0px">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/pusher_obstacle_standard.gif" width="100%" />
  </td>
<td style="border-top:none; border-bottom:none; padding:0px">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/pusher_obstacle_maxent.gif" width="100%" />
  </td>
</tr>
</table>
<h1 id="theory">Theory</h1>
<p>We now formally describe the technical results from the paper. The aim here is not to provide a full proof (see the paper Appendix for that), but instead to build some intuition for what the technical results say. Our main result is that, when you apply MaxEnt RL with some reward function and some dynamics, you are actually maximizing a lower bound on the robust RL objective. To explain this result, we must first define the MaxEnt RL objective: $J_{MaxEnt}(\pi; p, r)$ is the entropy-regularized cumulative return of policy $\pi$ when evaluated using dynamics $p(s’ \mid s, a)$ and reward function $r(s, a)$. While we will train the policy using one dynamics $p$, we will evaluate the policy on a different dynamics, $\tilde{p}(s’ \mid s, a)$, chosen by the adversary. We can now formally state our main result as follows:</p>
<p><script type="math/tex; mode=display">\min_{\tilde{p} \in \tilde{\mathcal{P}}(\pi)} J_\text{MaxEnt}(\pi; \tilde{p},
r) \ge \exp(J_\text{MaxEnt}(\pi; p, \bar{r}) + \log T.</script></p>
<p>The left-hand-side is the robust RL objective. It says that the adversary gets to choose whichever dynamics function $\tilde{p}(s’ \mid s, a)$ makes our policy perform as poorly as possible, subject to some constraints (as specified by the set $\tilde{\mathcal{P}}$). On the right-hand-side we have the MaxEnt RL objective (note that $\log T$ is a constant, and the function $\exp(\cdots)$ is always increasing). Thus, this objective says that a policy that has a high entropy-regularized reward (right hand-side) is guaranteed to also get high reward when evaluated on an adversarially-chosen dynamics.</p>
<p>The most important part of this equation is the set $\tilde{\mathcal{P}}$ of dynamics that the adversary can choose from. Our analysis describes precisely how this set is constructed and shows that, if we want a policy to be robust to a larger set of disturbances, all we have to do is increase the weight on the entropy term and decrease the weight on the reward term. Intuitively, the adversary must choose dynamics that are “close” to the dynamics on which the policy was trained. For example, in the special case where the dynamics are linear-Gaussian, this set corresponds to all perturbations where the original expected next state and the perturbed expected next state have a Euclidean distance less than $\epsilon$.</p>
<h1 id="more-experiments">More Experiments</h1>
<p>Our analysis predicts that MaxEnt RL should be robust to many types of disturbances. The first set of videos in this post showed that MaxEnt RL is robust to  static obstacles. MaxEnt RL is also robust to dynamic perturbations introduced in the middle of an episode. To demonstrate this, we took the same robotic pushing task and knocked the puck out of place in the middle of the episode. The videos below show that the policy learned by MaxEnt RL is more robust at handling these perturbations, as predicted by our analysis.</p>
<p style="text-align:center;">
<table style="width:70%; margin-left:auto; margin-right:auto">
<tr>
<td style="border-top:none; border-bottom:none; padding:0px">
<p style="text-align:center; margin-bottom:0px">Standard RL</p>
</td>
<td style="border-top:none; border-bottom:none; padding:0px">
<p style="text-align:center; margin-bottom:0px">MaxEnt RL</p>
</td>
</tr>
<tr>
<td style="border-top:none; border-bottom:none; padding:0px; vertical-align:middle;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/pusher_force_standard_v2_opt.gif" width="100%" />
  </td>
<td style="border-top:none; border-bottom:none; padding:0px; vertical-align:middle;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/pusher_force_maxent_v2_opt.gif" width="100%" />
  </td>
</tr>
</table>
<p style="text-align:center;"><i>The policy learned by MaxEntRL is robust to dynamic perturbations of the puck (red frames).<br />
</i></p>
</p>
<p>Our theoretical results suggest that, even if we optimize the environment perturbations so the agent does as poorly as possible, MaxEnt RL policies will still be robust. To demonstrate this capability, we trained both standard RL and MaxEnt RL on a peg insertion task shown below. During evaluation, we changed the position of the hole to try to make each policy fail. If we only moved the hole position a little bit ($\le$ 1 cm), both policies always solved the task. However, if we moved the hole position up to 2cm, the policy learned by standard RL almost never succeeded in inserting the peg, while the MaxEnt RL policy succeeded in 95% of trials. This experiment validates our theoretical findings that MaxEnt really is robust to (bounded) adversarial disturbances in the environment.</p>
<p style="text-align:center;">
<table>
<colgroup>
<col span="1" style="width: 27%;" />
<col span="1" style="width: 27%;" />
<col span="1" style="width: 45%;" />
</colgroup>
<tr>
<td style="border-top:none; border-bottom:none; padding:0px; vertical-align:middle;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/peg_standard_long.gif" width="100%" />
  </td>
<td style="border-top:none; border-bottom:none; padding:0px; vertical-align:middle;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/peg_maxent_long.gif" width="100%" />
  </td>
<td style="border-top:none; border-bottom:none; padding:0px; vertical-align:middle;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/maxent-robust-rl/peg_minimax_100.png" width="100%" />
  </td>
</tr>
<tr>
<td style="border-top:none; border-bottom:none; padding:0px">
<p style="text-align:center;">Standard RL</p>
</td>
<td style="border-top:none; border-bottom:none; padding:0px">
<p style="text-align:center;">MaxEnt RL</p>
</td>
<td style="border-top:none; border-bottom:none; padding:0px">
<p style="text-align:center;">Evaluation on adversarial perturbations</p>
</td>
</tr>
</table>
<p style="text-align:center;">
<i>MaxEnt RL is robust to adversarial perturbations of the hole (where the robot<br />
inserts the peg).</i></p>
</p>
<h1 id="conclusion">Conclusion</h1>
<p>In summary, <a href="https://arxiv.org/abs/2103.06257" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our paper</a> shows that a commonly-used type of RL algorithm, MaxEnt RL, is already solving a robust RL problem. We do not claim that MaxEnt RL will outperform purpose-designed robust RL algorithms. However, the striking simplicity of MaxEnt RL compared with other robust RL algorithms suggests that it may be an appealing alternative to practitioners hoping to equip their RL policies with an ounce of robustness.</p>
<p><strong>Acknowledgements</strong><br />
Thanks to Gokul Swamy, Diba Ghosh, Colin Li, and Sergey Levine for feedback on drafts of this post, and to Chloe Hsu and Daniel Seita for help with the blog.</p>
<hr />
<p>This post is based on the following paper:</p>
<ul>
<li><a href="https://arxiv.org/abs/2103.06257" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Maximum Entropy RL (Provably) Solves Some Robust RL Problems</a>. <br />
<a href="https://ben-eysenbach.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Benjamin Eysenbach</a> and <a href="https://people.eecs.berkeley.edu/~svlevine/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sergey Levine</a>.</li>
</ul>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Self-supervised policy adaptation during deployment</title>
		<link>https://robohub.org/self-supervised-policy-adaptation-during-deployment/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Fri, 26 Feb 2021 11:55:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<category><![CDATA[manipulation]]></category>
		<category><![CDATA[open source]]></category>
		<category><![CDATA[research]]></category>
		<guid isPermaLink="false">https://robohub.org/self-supervised-policy-adaptation-during-deployment/</guid>

					<description><![CDATA[<!-- twitter -->
<p>
<img src="https://bair.berkeley.edu/static/blog/ss-adaptation/0_header_0.gif" width="30%"><img src="https://bair.berkeley.edu/static/blog/ss-adaptation/0_header_1.gif" width="30%"><img src="https://bair.berkeley.edu/static/blog/ss-adaptation/0_header_2.gif" width="30%"><br><i>
Our method learns a task in a fixed, simulated environment and quickly adapts
to new environments (e.g. the real world) solely from online interaction during
deployment.
</i>
</p>

<p>The ability for humans to generalize their knowledge and experiences to new
situations is remarkable, yet poorly understood. For example, imagine a human
driver that has only ever driven around their city in clear weather. Even
though they never encountered true diversity in driving conditions, they have
acquired the fundamental skill of driving, and can adapt reasonably fast to
driving in neighboring cities, in rainy or windy weather, or even driving a
different car, without much practice nor additional driver&#8217;s lessons. While
humans excel at adaptation, building intelligent systems with common-sense
knowledge and the ability to quickly adapt to new situations is a long-standing
problem in artificial intelligence.</p>

<!--more-->

<p>
<img src="https://bair.berkeley.edu/static/blog/ss-adaptation/1_intro_0.png" width="80%"><br><i>
A robot trained to perform a given task in a lab environment may not generalize
to other environments, e.g. an environment with moving disco lights, even
though the task itself remains the same.
</i>
</p>

<p>In recent years, learning both perception and behavioral policies in an
end-to-end framework by deep Reinforcement Learning (RL) has been widely
successful, and has achieved impressive results such as superhuman performance
on Atari games played directly from screen pixels.  Although impressive, it has
become commonly understood that such policies fail to generalize to <em>even
subtle changes</em> in the environment - changes that humans are easily able to
adapt to. For this reason, RL has shown limited success beyond the game or
environment in which it was originally trained, which presents a significant
challenge in deployment of policies trained by RL in our diverse and
unstructured real world.</p>

<h1>Generalization by Randomization</h1>

<p>In applications of RL, practitioners have sought to improve the generalization
ability of policies by introducing randomization into the training environment
(e.g. a simulation), also known as <em>domain randomization</em>. By randomizing
elements of the training environment that are also expected to vary at
test-time, it is possible to learn policies that are <em>invariant</em> to certain
factors of variation. For autonomous driving, we may for example want our
policy to be robust to changes in lighting, weather, and road conditions, as
well as car models, nearby buildings, different city layouts, and so forth.
While the randomization quickly evolves into an elaborate engineering challenge
as more and more factors of variation are considered, the learning problem
itself also becomes harder, greatly decreasing the sample efficiency of
learning algorithms. It is therefore natural to ask: rather than learning a
policy robust to all conceivable environmental changes, can we instead <em>adapt</em>
a pre-trained policy to the new environment through interaction?</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/ss-adaptation/2_problem_0_optimized.gif" width="30%"><img src="https://bair.berkeley.edu/static/blog/ss-adaptation/2_problem_1_optimized.gif" width="30%"><br><i>
<b>Left</b>: training in a fixed environment. <b>Right</b>: training with
domain randomization.
</i>
</p>

<h1>Policy Adaptation</h1>

<p>A na&#239;ve way to adapt a policy to new environments is by fine-tuning parameters
using a reward signal. In real-world deployments, however, obtaining a reward
signal often requires human feedback or careful engineering, neither of which
are scalable solutions.</p>

<p>In recent work from our lab, we show that it is possible to adapt a pre-trained
policy to unseen environments, without any reward signal or human supervision.
A key insight is that, in the context of many deployments of RL, the
fundamental goal of the task remains the same, even though there may be a
mismatch in both visuals and underlying dynamics compared to the training
environment, e.g. a simulation. When training a policy in simulation and
deploying it in the real world (sim2real), there are often differences in
dynamics due to imperfections in the simulation, and visual inputs captured by
a camera are likely to differ from renderings of the simulation. Hence, the
source of these errors often lie in an imperfect world understanding rather
than misspecification of the task itself, and an agent&#8217;s interactions with a
new environment can therefore provide us with valuable information about the
disparity between its world understanding and reality.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/ss-adaptation/3_framework_0.png" width="48%"><img src="https://bair.berkeley.edu/static/blog/ss-adaptation/3_framework_1.png" width="48%"><br><i>
Illustration of our framework for adaptation. <b>Left</b>: training before
deployment. The RL objective is optimized together with a self-supervised
objective. <b>Right</b>: adaptation during deployment. We optimize only the
self-supervised objective, using observations collected through interaction
with the environment.
</i>
</p>

<p>To take advantage of this information we turn to the literature of
self-supervised learning. We propose <strong>PAD</strong>, a general framework for
adaptation of policies <em>during deployment</em>, by using self-supervision as a
proxy for the absent reward signal. A given policy network $\pi$ parameterized
by a collection of parameters $\theta$ is split sequentially into an encoder
$\pi_{e}$ and a policy head $\pi_{a}$ such that $a_{t} = \pi(s_{t}; \theta) =
\pi_{a}(\pi_{e}(s_{t}; \theta_{e}) ;\theta_{a})$ for a state $s_{t}$ and action
$a_{t}$ at time $t$. We then let $\pi_{s}$ be a self-supervised task head and
similarly let $\pi_{s}$ share the encoder $\pi_{e}$ with the policy head.
During training, we optimize a self-supervised objective jointly together with
the RL task, where the two tasks share part of a neural network. During
deployment, we can no longer assume access to a reward signal and are unable to
optimize the RL objective. However, we can still continue to optimize the
self-supervised objective using observations collected through interaction with
the new environment. At every step in the new environment, we update the policy
through self-supervision, using only the most recently collected observation:</p>

<p>where $L$ is a self-supervised objective. Assuming that gradients of the
self-supervised objective are sufficiently correlated with those of the RL
objective, any adaptation in the self-supervised task may also influence and
correct errors in the perception and decision-making of the policy.</p>

<p>In practice, we use an inverse dynamics model $a_{t} = \pi_{s}( \pi_e(s_{t}),
\pi_e(s_{t+1}))$, predicting the action taken in between two consecutive
observations. Because an inverse dynamics model connects observations directly
to actions, the policy can be adjusted for disparities both in visuals <em>and</em>
dynamics (e.g. lighting conditions or friction) between training and test
environments, solely through interaction with the new environment.</p>

<h1>Adapting policies to the real world</h1>

<p>We demonstrate the effectiveness of self-supervised policy adaptation (PAD) by
training policies for robotic manipulation tasks in simulation and adapting
them to the real world during deployment on a physical robot, taking
observations directly from an uncalibrated camera. We evaluate generalization
to a real robot environment that resembles the simulation, as well as two more
challenging settings: a table cloth with increased friction, and continuously
moving disco lights. In the demonstration below, we consider a Soft
Actor-Critic (SAC) agent trained with an Inverse Dynamics Model (IDM), with and
without the PAD adaptation mechanism.</p>

&#60;!--
<div class="videoWrapper">
  
</div>

<p style="text-align:center">
<i>
Transferring a policy from simulation to the real world. <b>SAC+IDM</b> is a
Soft Actor-Critic (SAC) policy trained with an Inverse Dynamics Model (IDM),
and <b>SAC+IDM (PAD)</b> is the same policy but with the addition of policy
adaptation during deployment on the robot.
</i>
</p>
--&#62;

<p>
<img src="https://bair.berkeley.edu/static/blog/ss-adaptation/4_comparison_0.gif" width="100%"><br><i>
Transferring a policy from simulation to the real world. <b>SAC+IDM</b> is a
Soft Actor-Critic (SAC) policy trained with an Inverse Dynamics Model (IDM),
and <b>SAC+IDM (PAD)</b> is the same policy but with the addition of policy
adaptation during deployment on the robot.
</i>
</p>

<p>PAD adapts to changes in both visuals and dynamics, and nearly recovers the
original success rate of the simulated environment. Policy adaptation is
especially effective when the test environment differs from the training
environment in multiple ways, e.g. where both visuals <em>and</em> physical properties
such as object dimensionality and friction differ. Because it is often
difficult to formally specify the elements that vary between a simulation and
the real world, policy adaptation may be a promising alternative to domain
randomization techniques in such settings.</p>

<h1>Benchmarking generalization</h1>

<p>Simulations provide a good platform for more comprehensive evaluation of RL
algorithms. Together with PAD, we release <a href="https://github.com/nicklashansen/dmcontrol-generalization-benchmark" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DMControl Generalization
Benchmark</a>, a new benchmark for generalization in RL based on the <em>DeepMind
Control Suite</em>, a popular benchmark for continuous control from images. In the
DMControl Generalization Benchmark, agents are trained in a fixed environment
and deployed in new environments with e.g. randomized colors or continuously
changing video backgrounds. We consider an SAC agent trained with an IDM, with
and without adaptation, and compare to CURL, a contrastive method discussed in
<a href="https://bair.berkeley.edu/blog/2020/07/19/curl-rad/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">a previous post</a>. We compare the generalization ability of methods in the
visualization below, and generally find that PAD can adapt even in
non-stationary environments, a challenging problem setting where non-adaptive
methods tend to fail. While CURL is found to generalize no better than the
non-adaptive SAC trained with an IDM, agents can still benefit from the
training signal that CURL provides during the training phase. Algorithms that
learn both during training and deployment, and from multiple training signals,
may therefore be preferred.</p>

&#60;!--
<div class="videoWrapper">
  
</div>

<p style="text-align:center">
<i>
-- -- --  SAC+IDM        --  CURL    &#8212; &#8212; &#8212;  SAC+IDM (PAD)
</i>
</p>
--&#62;

<p>
<img src="https://bair.berkeley.edu/static/blog/ss-adaptation/4_comparison_1.gif" width="100%"><br><i>
Generalization to an environment with video background. <b>CURL</b> is a
contrastive method, <b>SAC+IDM</b> is a Soft Actor-Critic (SAC) policy trained
with an Inverse Dynamics Model (IDM), and <b>SAC+IDM (PAD)</b> is the same
policy but with the addition of policy adaptation during deployment.
</i>
</p>

<h1>Summary</h1>

<p>Previous work addresses the problem of generalization in RL by randomization,
which requires anticipation of environmental changes and is known to not scale
well. We formulate an alternative problem setting in vision-based RL: can we
instead <em>adapt</em> a pre-trained policy to unseen environments, without any
rewards or human feedback? We find that adapting policies through a
self-supervised objective - solely from interactions in the new environment -
is a promising alternative to domain randomization when the target environment
is truly unknown. In the future, we ultimately envision agents that
continuously learn and adapt to their surroundings, and are capable of learning
both from explicit human feedback <em>and</em> through unsupervised interaction with
the environment.</p>

<p>This post is based on the following paper:</p>

<ul><li><strong>Self-Supervised Policy Adaptation during Deployment</strong><br><strong>Nicklas Hansen</strong>, Rishabh Jangir, Yu Sun, Guillem Aleny&#225;, Pieter Abbeel, Alexei A. Efros, Lerrel Pinto, <strong>Xiaolong Wang</strong><br>
Ninth International Conference on Learning Representations (ICLR), 2021<br><a href="https://arxiv.org/abs/2007.04309" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">arXiv</a>, <a href="https://nicklashansen.github.io/PAD/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project Website</a>, <a href="https://github.com/nicklashansen/policy-adaptation-during-deployment" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Code</a></li>
</ul>]]></description>
										<content:encoded><![CDATA[<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/ss-adaptation/0_header_0.gif" width="30%" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/ss-adaptation/0_header_1.gif" width="30%" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/ss-adaptation/0_header_2.gif" width="30%" /><br />
<i><br />
Our method learns a task in a fixed, simulated environment and quickly adapts<br />
to new environments (e.g. the real world) solely from online interaction during<br />
deployment.<br />
</i>
</p>
<p>The ability for humans to generalize their knowledge and experiences to new situations is remarkable, yet poorly understood. For example, imagine a human driver that has only ever driven around their city in clear weather. Even though they never encountered true diversity in driving conditions, they have acquired the fundamental skill of driving, and can adapt reasonably fast to driving in neighboring cities, in rainy or windy weather, or even driving a different car, without much practice nor additional driver’s lessons. While humans excel at adaptation, building intelligent systems with common-sense knowledge and the ability to quickly adapt to new situations is a long-standing problem in artificial intelligence.<br />
<span id="more-199179"></span></p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/ss-adaptation/1_intro_0.png" width="80%" /><br />
<i><br />
A robot trained to perform a given task in a lab environment may not generalize<br />
to other environments, e.g. an environment with moving disco lights, even<br />
though the task itself remains the same.<br />
</i>
</p>
<p>In recent years, learning both perception and behavioral policies in an end-to-end framework by deep Reinforcement Learning (RL) has been widely successful, and has achieved impressive results such as superhuman performance on Atari games played directly from screen pixels. Although impressive, it has become commonly understood that such policies fail to generalize to even subtle changes in the environment &#8211; changes that humans are easily able to adapt to. For this reason, RL has shown limited success beyond the game or environment in which it was originally trained, which presents a significant challenge in deployment of policies trained by RL in our diverse and unstructured real world.</p>
<h1 id="generalization-by-randomization">Generalization by Randomization</h1>
<p>In applications of RL, practitioners have sought to improve the generalization ability of policies by introducing randomization into the training environment (e.g. a simulation), also known as domain randomization. By randomizing elements of the training environment that are also expected to vary at test-time, it is possible to learn policies that are invariant to certain factors of variation. For autonomous driving, we may for example want our policy to be robust to changes in lighting, weather, and road conditions, as well as car models, nearby buildings, different city layouts, and so forth. While the randomization quickly evolves into an elaborate engineering challenge as more and more factors of variation are considered, the learning problem itself also becomes harder, greatly decreasing the sample efficiency of learning algorithms. It is therefore natural to ask: rather than learning a policy robust to all conceivable environmental changes, can we instead adapt a pre-trained policy to the new environment through interaction?</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/ss-adaptation/2_problem_0_optimized.gif" width="30%" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/ss-adaptation/2_problem_1_optimized.gif" width="30%" /><br />
<i><br />
<b>Left</b>: training in a fixed environment. <b>Right</b>: training with<br />
domain randomization.<br />
</i>
</p>
<h1 id="policy-adaptation">Policy Adaptation</h1>
<p>A naïve way to adapt a policy to new environments is by fine-tuning parameters using a reward signal. In real-world deployments, however, obtaining a reward signal often requires human feedback or careful engineering, neither of which are scalable solutions.</p>
<p>In recent work from our lab, we show that it is possible to adapt a pre-trained policy to unseen environments, without any reward signal or human supervision. A key insight is that, in the context of many deployments of RL, the fundamental goal of the task remains the same, even though there may be a mismatch in both visuals and underlying dynamics compared to the training environment, e.g. a simulation. When training a policy in simulation and deploying it in the real world (sim2real), there are often differences in dynamics due to imperfections in the simulation, and visual inputs captured by a camera are likely to differ from renderings of the simulation. Hence, the source of these errors often lie in an imperfect world understanding rather than misspecification of the task itself, and an agent’s interactions with a new environment can therefore provide us with valuable information about the disparity between its world understanding and reality.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/ss-adaptation/3_framework_0.png" width="48%" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/ss-adaptation/3_framework_1.png" width="48%" /><br />
<br />
<i><br />
Illustration of our framework for adaptation. <b>Left</b>: training before<br />
deployment. The RL objective is optimized together with a self-supervised<br />
objective. <b>Right</b>: adaptation during deployment. We optimize only the<br />
self-supervised objective, using observations collected through interaction<br />
with the environment.<br />
</i>
</p>
<p>To take advantage of this information we turn to the literature of self-supervised learning. We propose <strong>PAD</strong>, a general framework for adaptation of policies <em>during deployment</em>, by using self-supervision as a proxy for the absent reward signal. A given policy network $\pi$ parameterized by a collection of parameters $\theta$ is split sequentially into an encoder $\pi_{e}$ and a policy head $\pi_{a}$ such that $a_{t} = \pi(s_{t}; \theta) = \pi_{a} (\pi_{e}(s_{t}; \theta_{e}) ;\theta_{a})$ for a state $s_{t}$ and action $a_{t}$ at time $t$. We then let $\pi_{s}$ be a self-supervised task head and similarly let $\pi_{s}$ share the encoder $\pi_{e}$ with the policy head. During training, we optimize a self-supervised objective jointly together with the RL task, where the two tasks share part of a neural network. During deployment, we can no longer assume access to a reward signal and are unable to optimize the RL objective. However, we can still continue to optimize the self-supervised objective using observations collected through interaction with the new environment. At every step in the new environment, we update the policy through self-supervision, using only the most recently collected observation:</p>
<p>$$s_t \sim p(s_t | a_{t-1}, s_{t-1}) \\<br />
\theta_{e}(t) = \theta_{e}(t-1) &#8211; \nabla_{\theta_{e}}L(s_{t}; \theta_{s}(t-1), \theta_{e}(t-1))$$</p>
<p>where L is a self-supervised objective. Assuming that gradients of the self-supervised objective are sufficiently correlated with those of the RL objective, any adaptation in the self-supervised task may also influence and correct errors in the perception and decision-making of the policy.</p>
<p>In practice, we use an inverse dynamics model $a_{t} = \pi_{s}( \pi_e(s_{t}), \pi_e(s_{t+1}))$, predicting the action taken in between two consecutive observations. Because an inverse dynamics model connects observations directly to actions, the policy can be adjusted for disparities both in visuals and dynamics (e.g. lighting conditions or friction) between training and test environments, solely through interaction with the new environment.</p>
<h1 id="adapting-policies-to-the-real-world">Adapting policies to the real world</h1>
<p>We demonstrate the effectiveness of self-supervised policy adaptation (PAD) by training policies for robotic manipulation tasks in simulation and adapting them to the real world during deployment on a physical robot, taking observations directly from an uncalibrated camera. We evaluate generalization to a real robot environment that resembles the simulation, as well as two more challenging settings: a table cloth with increased friction, and continuously moving disco lights. In the demonstration below, we consider a Soft Actor-Critic (SAC) agent trained with an Inverse Dynamics Model (IDM), with and without the PAD adaptation mechanism.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/ss-adaptation/4_comparison_0.gif" width="100%" /><br />
<i><br />
Transferring a policy from simulation to the real world. <b>SAC+IDM</b> is a<br />
Soft Actor-Critic (SAC) policy trained with an Inverse Dynamics Model (IDM),<br />
and <b>SAC+IDM (PAD)</b> is the same policy but with the addition of policy<br />
adaptation during deployment on the robot.<br />
</i>
</p>
<p>PAD adapts to changes in both visuals and dynamics, and nearly recovers the original success rate of the simulated environment. Policy adaptation is especially effective when the test environment differs from the training environment in multiple ways, e.g. where both visuals and physical properties such as object dimensionality and friction differ. Because it is often difficult to formally specify the elements that vary between a simulation and the real world, policy adaptation may be a promising alternative to domain randomization techniques in such settings.</p>
<h1 id="benchmarking-generalization">Benchmarking generalization</h1>
<p>Simulations provide a good platform for more comprehensive evaluation of RL algorithms. Together with PAD, we release <a href="https://github.com/nicklashansen/dmcontrol-generalization-benchmark" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DMControl Generalization Benchmark</a>, a new benchmark for generalization in RL based on the DeepMind Control Suite, a popular benchmark for continuous control from images. In the DMControl Generalization Benchmark, agents are trained in a fixed environment and deployed in new environments with e.g. randomized colors or continuously changing video backgrounds. We consider an SAC agent trained with an IDM, with and without adaptation, and compare to CURL, a contrastive method discussed in <a href="https://bair.berkeley.edu/blog/2020/07/19/curl-rad/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">a previous post</a>. We compare the generalization ability of methods in the visualization below, and generally find that PAD can adapt even in non-stationary environments, a challenging problem setting where non-adaptive methods tend to fail. While CURL is found to generalize no better than the non-adaptive SAC trained with an IDM, agents can still benefit from the training signal that CURL provides during the training phase. Algorithms that learn both during training and deployment, and from multiple training signals, may therefore be preferred.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/ss-adaptation/4_comparison_1.gif" width="100%" /><br />
<i><br />
Generalization to an environment with video background. <b>CURL</b> is a<br />
contrastive method, <b>SAC+IDM</b> is a Soft Actor-Critic (SAC) policy trained<br />
with an Inverse Dynamics Model (IDM), and <b>SAC+IDM (PAD)</b> is the same<br />
policy but with the addition of policy adaptation during deployment.<br />
</i>
</p>
<h1 id="summary">Summary</h1>
<p>Previous work addresses the problem of generalization in RL by randomization, which requires anticipation of environmental changes and is known to not scale well. We formulate an alternative problem setting in vision-based RL: can we instead adapt a pre-trained policy to unseen environments, without any rewards or human feedback? We find that adapting policies through a self-supervised objective &#8211; solely from interactions in the new environment &#8211; is a promising alternative to domain randomization when the target environment is truly unknown. In the future, we ultimately envision agents that continuously learn and adapt to their surroundings, and are capable of learning both from explicit human feedback and through unsupervised interaction with the environment.</p>
<p>This post is based on the following paper:</p>
<ul>
<li><strong>Self-Supervised Policy Adaptation during Deployment</strong><br />
<strong>Nicklas Hansen</strong>, Rishabh Jangir, Yu Sun, Guillem Alenyá, Pieter Abbeel, Alexei A. Efros, Lerrel Pinto, <strong>Xiaolong Wang</strong><br />
Ninth International Conference on Learning Representations (ICLR), 2021<br />
<a href="https://arxiv.org/abs/2007.04309" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">arXiv</a>, <a href="https://nicklashansen.github.io/PAD/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project Website</a>, <a href="https://github.com/nicklashansen/policy-adaptation-during-deployment" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Code</a></li>
</ul>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Plan2Explore: Active model-building for self-supervised visual reinforcement learning</title>
		<link>https://robohub.org/plan2explore-active-model-building-for-self-supervised-visual-reinforcement-learning/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Tue, 06 Oct 2020 06:01:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/plan2explore-active-model-building-for-self-supervised-visual-reinforcement-learning/</guid>

					<description><![CDATA[To operate successfully in unstructured open-world environments, autonomous
intelligent agents need to solve many different tasks and learn new tasks
quickly. Reinforcement learning has enabled artificial agents to solve complex
tasks both in sim...]]></description>
										<content:encoded><![CDATA[<p><strong>By <a href="https://www.seas.upenn.edu/~oleh/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Oleh Rybkin</a>, <a href="https://danijar.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Danijar Hafner</a> and <a href="https://www.cs.cmu.edu/~dpathak/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Deepak Pathak</a></strong></p>
<p>To operate successfully in unstructured open-world environments, autonomous intelligent agents need to solve many different tasks and learn new tasks quickly. Reinforcement learning has enabled artificial agents to solve complex tasks both in <a href="https://deepmind.com/research/case-studies/alphago-the-story-so-far" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">simulation</a> and <a href="https://ai.googleblog.com/2018/06/scalable-deep-reinforcement-learning.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">real-world</a>. However, it requires collecting large amounts of experience in the environment for each individual task. </p>
<p><span id="more-195754"></span></p>
<p>Self-supervised reinforcement learning has emerged <a href="https://pathak22.github.io/noreward-rl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">as</a> <a href="https://arxiv.org/abs/1903.03698" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">an</a>  <a href="https://arxiv.org/abs/1907.01657" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">alternative</a>, where the agent only follows an intrinsic objective that is independent of any individual task, analogously to <a href="https://youtu.be/VsnQf7exv5I" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">unsupervised representation learning</a>. After acquiring general and reusable knowledge about the environment through self-supervision, the agent can adapt to specific downstream tasks more efficiently.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/plan2explore/figure1_teaser.gif" height="" width="90%" /><br />

</p>
<p>In this post, we explain our recent publication that develops <a href="https://ramanans1.github.io/plan2explore/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Plan2Explore</a>. While many recent papers on self-supervised reinforcement learning have focused on <a href="https://spinningup.openai.com/en/latest/spinningup/rl_intro2.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">model-free</a> agents, our agent learns an internal <a href="https://bair.berkeley.edu/blog/2019/12/12/mbpo/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">world model</a>  that predicts the future outcomes of potential actions. The world model captures general knowledge, allowing Plan2Explore to quickly solve new tasks through planning in its own imagination. The world model further enables the agent to explore what it expects to be novel, rather than repeating what it found novel in the past. Plan2Explore obtains state-of-the-art zero-shot and few-shot performance on continuous control benchmarks with high-dimensional input images. To make it easy to experiment with our agent, we are open-sourcing the complete <a href="https://github.com/ramanans1/plan2explore" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">source code</a>.</p>
<h1 id="how-does-plan2explore-work">How does Plan2Explore work?</h1>
<p>At a high level, Plan2Explore works by training a world model, exploring to maximize the information gain for the world model, and using the world model at test time to solve new tasks (see figure above). Thanks to effective exploration, the learned world model is general and captures information that can be used to solve multiple new tasks with no or few additional environment interactions. We discuss each part of the Plan2Explore algorithm individually below. We assume a basic understanding of reinforcement learning in this post and otherwise recommend <a href="https://spinningup.openai.com/en/latest/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">these</a> <a href="http://rail.eecs.berkeley.edu/deeprlcourse/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">materials</a>  as an introduction.</p>
<h1 id="learning-the-world-model">Learning the world model</h1>
<p>Plan2Explore learns a world model that predicts future outcomes given past observations <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-eebada2fc0fd8eebc217dc268a0775a6_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#111;&#95;&#123;&#49;&#58;&#116;&#125;" title="Rendered by QuickLaTeX.com" height="11" width="25" style="vertical-align: -3px;"/> and actions <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-4b1b487c9ee6c18e18ceb35b212f52fc_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#97;&#95;&#123;&#49;&#58;&#116;&#125;" title="Rendered by QuickLaTeX.com" height="11" width="25" style="vertical-align: -3px;"/> (see figure below). To handle high-dimensional image observations, we encode them into lower-dimensional features <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-14b463d0ecd5b350ced6cf1d6a12eef3_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#104;" title="Rendered by QuickLaTeX.com" height="12" width="10" style="vertical-align: 0px;"/> and use an <a href="https://ai.googleblog.com/2019/02/introducing-planet-deep-planning.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RSSM</a> model that predicts forward in a compact latent state-space <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ae1901659f469e6be883797bfd30f4f8_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#115;" title="Rendered by QuickLaTeX.com" height="8" width="8" style="vertical-align: 0px;"/>, from which the observations can be decoded. The latent state aggregates information from past observations that is helpful for future prediction, and is learned end-to-end using a variational objective.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/plan2explore/figure2_model.gif" height="" width="90%" /><br />

</p>
<h1 id="a-novelty-metric-for-active-model-building">A novelty metric for active model-building</h1>
<p>To learn an accurate and general world model we need an exploration strategy that collects new and informative data. To achieve this, Plan2Explore uses a novelty metric derived from the model itself. The novelty metric measures the expected information gained about the environment upon observing the new data. As the figure below shows, this is approximated by the disagreement <a href="https://arxiv.org/abs/1612.01474" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">of</a> <a href="https://pathak22.github.io/exploration-by-disagreement/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">an</a>  <a href="https://arxiv.org/abs/2002.08791" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ensemble</a>  of <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ea9c87a513e4a72624155d392fae86e2_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#75;" title="Rendered by QuickLaTeX.com" height="12" width="16" style="vertical-align: 0px;"/> latent models. Intuitively, large latent disagreement reflects high model uncertainty, and obtaining the data point would reduce this uncertainty. By maximizing latent disagreement, Plan2Explore selects actions that lead to the largest information gain, therefore improving the model as quickly as possible.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/plan2explore/figure3_disagreement.gif" height="" width="65%" /><br />

</p>
<h1 id="planning-for-future-novelty">Planning for future novelty</h1>
<p>To effectively maximize novelty, we need to know which parts of the environment are still unexplored. Most prior work on self-supervised exploration used model-free methods that reinforce past behavior that resulted in novel experience. This makes these methods slow to explore: since they can only repeat exploration behavior that was successful in the past, they are unlikely to stumble onto something novel. In contrast, Plan2Explore plans for expected novelty by measuring model uncertainty of imagined future outcomes. By seeking trajectories that have the highest uncertainty, Plan2Explore explores exactly the parts of the environments that were previously unknown.</p>
<p>To choose actions <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-5c53d6ebabdbcfa4e107550ea60b1b19_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#97;" title="Rendered by QuickLaTeX.com" height="8" width="9" style="vertical-align: 0px;"/> that optimize the exploration objective, Plan2Explore leverages the learned world model as shown in the figure below. The actions are selected to maximize the expected novelty of the entire future sequence <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-e8cd339428791a82c2e1ace590edf5ba_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#115;&#95;&#123;&#116;&#58;&#84;&#125;" title="Rendered by QuickLaTeX.com" height="11" width="27" style="vertical-align: -3px;"/>, using imaginary rollouts of the world model to estimate the novelty. To solve this optimization problem, we use the <a href="https://ai.googleblog.com/2020/03/introducing-dreamer-scalable.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Dreamer</a> agent, which learns a policy <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-aa37ac11527640d479cf10422b69c87e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#112;&#105;&#95;&#92;&#112;&#104;&#105;" title="Rendered by QuickLaTeX.com" height="14" width="18" style="vertical-align: -6px;"/> using a value function and analytic gradients through the model. The policy is learned completely inside the imagination of the world model. During exploration, this imagination training ensures that our exploration policy is always up-to-date with the current world model and collects data that are still novel.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/plan2explore/figure4_policy.gif" height="" width="90%" /><br />

</p>
<h1 id="curiosity-driven-exploration-behavior">Curiosity-driven exploration behavior</h1>
<p>We evaluate Plan2Explore on 20 continuous control tasks from the <a href="https://github.com/deepmind/dm_control" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DeepMind Control Suite</a>. The agent only has access to image observations and no proprioceptive information.  Instead of random exploration, which fails to take the agent far from the initial position, Plan2Explore leads to diverse movement strategies like jumping, running, and flipping. Later, we will see that these are effective practice episodes that enable the agent to quickly learn to solve various continuous control tasks.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/plan2explore/figure5_gif1.gif" height="190" width="" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/plan2explore/figure5_gif2.gif" height="190" width="" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/plan2explore/figure5_gif3.gif" height="190" width="" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/plan2explore/figure5_gif4.gif" height="190" width="" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/plan2explore/figure5_gif5.gif" height="190" width="" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/plan2explore/figure5_gif6.gif" height="190" width="" /><br />

</p>
<h1 id="solving-tasks-with-the-world-model">Solving tasks with the world model</h1>
<p>Once an accurate and general world model is learned, we test Plan2Explore on previously unseen tasks. Given a task specified with a reward function, we use the model to optimize a policy for that task. Similar to our exploration procedure, we optimize a new value function and a new policy head for the downstream task. This optimization uses only predictions imagined by the model, enabling Plan2Explore to solve new downstream tasks in a zero-shot manner without any additional interaction with the world.</p>
<p>The following plot shows the performance of Plan2Explore on tasks from DM Control Suite. Before 1 million environment steps, the agent doesn’t know the task and simply explores. The agent solves the task as soon as it is provided at 1 million steps, and keeps improving fast in a few-shot regime after that.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/plan2explore/figure6_plot.png" height="" width="" /><br />

</p>
<p>Plan2Explore (<font color="green"><strong>—</strong></font>) is able to solve most of the tasks we benchmarked. Since prior work on self-supervised reinforcement learning used model-free agents that are not able to adapt in a zero-shot manner (<a href="https://pathak22.github.io/noreward-rl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ICM</a>, <font color="blue"><strong>—</strong></font>), or did not use image observations, we compare by adapting this prior work to our model-based plan2explore setup. Our latent disagreement objective outperforms other previously proposed objectives. More interestingly, the final performance of Plan2Explore is comparable to the state-of-the-art  <a href="https://ai.googleblog.com/2020/03/introducing-dreamer-scalable.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">oracle</a> agent that requires task rewards throughout training (<font color="yellow"><strong>—</strong></font>). In our <a href="https://arxiv.org/abs/2005.05960" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a>, we further report performance of Plan2Explore in the zero-shot setting where the agent needs to solve the task before any task-oriented practice.</p>
<h1 id="future-directions">Future directions</h1>
<p>Plan2Explore demonstrates that effective behavior can be learned through self-supervised exploration only. This opens multiple avenues for future research:</p>
<ul>
<li>
<p>First, to apply self-supervised RL to a variety of settings, future work will investigate different ways of specifying the task and deriving behavior from the world model. For example, the task could be specified with a demonstration, description of the desired goal state, or communicated to the agent in natural language.</p>
</li>
<li>
<p>Second, while Plan2Explore is completely self-supervised, in many cases a weak supervision signal is available, such as in hard exploration games, human-in-the-loop learning, or real life. In such a semi-supervised setting, it is interesting to investigate how weak supervision can be used to steer exploration towards the relevant parts of the environment.</p>
</li>
<li>
<p>Finally, Plan2Explore has the potential to improve the data efficiency of real-world robotic systems, where exploration is costly and time-consuming, and the final task is often unknown in advance.</p>
</li>
</ul>
<p>By designing a scalable way of planning to explore in unstructured environments with visual observations, Plan2Explore provides an important step toward self-supervised intelligent machines.</p>
<hr />
<p>We would like to thank Georgios Georgakis for the useful feedback.</p>
<p>This post is based on the following paper:</p>
<ul class="list-external-links">
<li>Planning to Explore via Self-Supervised World Models<br /> Ramanan Sekar*, Oleh Rybkin*, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, Deepak Pathak<br /> Thirty-seventh International Conference Machine Learning (ICML), 2020.<br /> <a href="https://arxiv.org/abs/2005.05960" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">arXiv</a>, <a href="https://ramanans1.github.io/plan2explore/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project Website</a></li>
</ul>
<p>This article was initially published on the <a href="http://bair.berkeley.edu/blog/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>AWAC: Accelerating online reinforcement learning with offline datasets</title>
		<link>https://robohub.org/awac-accelerating-online-reinforcement-learning-with-offline-datasets/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 30 Sep 2020 14:22:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/awac-accelerating-online-reinforcement-learning-with-offline-datasets/</guid>

					<description><![CDATA[<p>
<img src="https://bair.berkeley.edu/static/blog/awac/01_hand.gif" height="190" width=""><img src="https://bair.berkeley.edu/static/blog/awac/02_hand.gif" height="190" width=""><img src="https://bair.berkeley.edu/static/blog/awac/03_hand.gif" height="190" width=""><img src="https://bair.berkeley.edu/static/blog/awac/04_hand.gif" height="190" width=""><br><i>Our method learns complex behaviors by training offline from prior datasets
(expert demonstrations, data from previous experiments, or random exploration
data) and then fine-tuning quickly with online interaction.</i>
</p>

<p>Robots trained with reinforcement learning (RL) have the potential to be used
across a huge variety of challenging real world problems. To apply RL to a new
problem, you typically set up the environment, define a reward function, and
train the robot to solve the task by allowing it to explore the new environment
from scratch. While this may eventually work, these &#8220;online&#8221; RL methods are
data hungry and repeating this data inefficient process for every new problem
makes it difficult to apply online RL to real world robotics problems. What if
instead of repeating the data collection and learning process from scratch
every time, we were able to reuse data across multiple problems or experiments?
By doing so, we could greatly reduce the burden of data collection with every
new problem that is encountered. With hundreds to thousands of robot
experiments being constantly run, it is of crucial importance to devise an RL
paradigm that can effectively use the large amount of already available data
while still continuing to improve behavior on new tasks.</p>

<p>The first step towards moving RL towards a data driven paradigm is to consider
the general idea of offline (batch) RL. Offline RL considers the problem of
learning optimal policies from arbitrary off-policy data, without any further
exploration. This is able to eliminate the data collection problem in RL, and
incorporate data from arbitrary sources including other robots or
teleoperation. However, depending on the quality of available data and the
problem being tackled, we will often need to augment offline training with
targeted online improvement. This problem setting actually has unique
challenges of its own. In this blog post, we discuss how we can move RL from
training from scratch with every new problem to a paradigm which is able to
reuse prior data effectively, with some offline training followed by online
finetuning.</p>

<!--more-->

<p>
<img src="https://bair.berkeley.edu/static/blog/awac/05_fig1.png" width="90%"><br><i>Figure 1: The problem of accelerating online RL with offline datasets. In
(1), the robot learns a policy entirely from an offline dataset. In (2), the
robot gets to interact with the world and collect on-policy samples to improve
the policy beyond what it could learn offline.</i>
</p>

<h1>Challenges in Offline RL with Online Fine-tuning</h1>

<p>We analyze the challenges in the problem of learning from offline data and
subsequent fine-tuning, using the standard benchmark HalfCheetah locomotion
task. The following experiments are conducted with a prior dataset consisting
of 15 demonstrations from an expert policy and 100 suboptimal trajectories
sampled from a behavioral clone of these demonstrations.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/awac/06_fig2.png" height="250"><img src="https://bair.berkeley.edu/static/blog/awac/07_fig2.png" height="250"><br><i>Figure 2: On-policy methods are slow to learn compared to off-policy
methods, due to the ability of off-policy methods to &#8220;stitch" good trajectories
together, illustrated on the left. Right: in practice, we see slow online
improvement using on-policy methods.</i>
</p>

<h2>1. Data Efficiency</h2>

<p>A simple way to utilize prior data such as demonstrations for RL is to
pre-train a policy with imitation learning, and fine-tune with on-policy RL
algorithms such as <a href="https://arxiv.org/abs/1910.00177" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">AWR</a> or <a href="https://arxiv.org/abs/1709.10087" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DAPG</a>.
This has two drawbacks.  First, the prior data may not be optimal so imitation
learning may be ineffective. Second, on-policy fine-tuning is data inefficient
as it does not reuse the prior data in the RL stage. For real-world robotics,
data efficiency is vital. Consider the robot on the right, trying to reach the
goal state with prior trajectory $\tau_1$ and $\tau_2$. On-policy methods
cannot effectively use this data, but off-policy algorithms that do dynamic
programming can, by effectively &#8220;stitching&#8221; $\tau_1$ and $\tau_2$ together with
the use of a value function or model. This effect can be seen in the learning
curves in Figure 2, where on-policy methods are an order of magnitude slower
than off-policy actor-critic methods.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/awac/08_fig3.png" height="220"><img src="https://bair.berkeley.edu/static/blog/awac/09_fig3.png" height="220"><img src="https://bair.berkeley.edu/static/blog/awac/10_fig3.png" height="220"><br><i>Figure 3: Bootstrapping error is an issue when using off-policy RL for
offline training. Left: an erroneous Q value far away from the data is
exploited by the policy, resulting in a poor update of the Q function. Middle:
as a result, the robot may take actions that are out of distribution. Right:
bootstrap error causes poor offline pretraining when using SAC and its
variants.</i>
</p>

<h2>2. Bootstrapping Error</h2>

<p>Actor-critic methods can in principle learn efficiently from off-policy data by
estimating a value estimate $V(s)$ or action-value estimate $Q(s, a)$ of future
returns by Bellman bootstrapping. However, when standard off-policy
actor-critic methods are applied to our problem (we use <a href="https://arxiv.org/abs/1801.01290" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SAC</a>), they perform poorly, as shown
in Figure 3: despite having a prior dataset in the replay buffer, these
algorithms do not benefit significantly from offline training (as seen by the
comparison between the SAC(scratch) and SACfD(prior) lines in Figure 3).
Moreover, even if the policy is pre-trained by behavior cloning (&#8220;SACfD
(pretrain)&#8221;) we still observe an initial decrease in performance.</p>

<p>This challenge can be attributed to off-policy bootstrapping error
accumulation. During training, the Q estimates will not be fully accurate,
particularly in extrapolating actions that are not present in the data. The
policy update exploits overestimated Q values, making the estimated Q values
worse. The issue is illustrated in the figure: incorrect Q values result in an
incorrect update to the target Q values, which may result in the robot taking a
poor action.</p>

<h2>3. Non-stationary Behavior Models</h2>

<p>Prior offline RL algorithms such as <a href="https://arxiv.org/pdf/1812.02900.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BCQ</a>, <a href="https://arxiv.org/abs/1906.00949" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BEAR</a>, and <a href="https://arxiv.org/abs/1911.11361" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BRAC</a> propose to address the
bootstrapping issue by preventing the policy from straying too far from the
data. The key idea is to prevent bootstrapping error by constraining the policy
$\pi$ close to the &#8220;behavior policy&#8221; $\pi_\beta$: the actions that are present
in the replay buffer. The idea is illustrated in the figure below: by sampling
actions from $\pi_\beta$, you avoid exploiting incorrect Q values far away from
the data distribution.</p>

<p><img src="https://bair.berkeley.edu/static/blog/awac/11_q.png" height="250" hspace="40" align="right"></p>

<p>However, $\pi_\beta$ is typically not known, especially for offline data, and
must be estimated from the data itself. Many offline RL algorithms (BEAR, BCQ,
<a href="https://arxiv.org/abs/2002.08396" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ABM</a>) explicitly fit a parametric
model to samples from the replay buffer for the distribution $\pi_\beta$. After
forming an estimate $\hat{\pi}_\beta$, prior methods implement the policy
constraint in various ways, including penalties on the policy update (BEAR,
BRAC) or architecture choices for sampling actions for policy training (BCQ,
ABM).</p>

<p>While offline RL algorithms with constraints perform well offline, they
struggle to improve with fine-tuning, as shown in the third plot in Figure 1.
We see that the purely offline RL performance (at &#8220;0K&#8221; in Fig.1) is much
better than SAC. However, with additional iterations of online fine-tuning, the
performance increases very slowly (as seen from the slope of the BEAR curve in
Fig 1). What causes this phenomenon?</p>

<p>The issue is in fitting an accurate behavior model as data is collected online
during fine-tuning. In the offline setting, behavior models must only be
trained once, but in the online setting, the behavior model must be updated
online to track incoming data. Training density models online (in the
&#8220;streaming&#8221; setting) is a challenging research problem, made more difficult by
a potentially complex multi-modal behavior distribution induced by the mixture
of online and offline data. In order to address our problem setting, we require
an off-policy RL algorithm that constrains the policy to prevent offline
instability and error accumulation, but is not so conservative that it prevents
online fine-tuning due to imperfect behavior modeling. Our proposed algorithm,
which we discuss in the next section, accomplishes this by employing an
implicit constraint, which does not require any explicit modeling of the
behavior policy.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/awac/12_fig4.png" height="300"><img src="https://bair.berkeley.edu/static/blog/awac/13_fig4.png" height="300"><br><i>Figure 4: an illustration of AWAC. High-advantage transitions are regressed
on with high weight, while low advantage transitions have low weight. Right:
algorithm pseudocode.</i>
</p>

<h1>Advantage Weighted Actor Critic</h1>

<p>In order to avoid these issues, we propose an extremely simple algorithm -
advantage weighted actor critic (AWAC).  AWAC avoids the pitfalls in the
previous section with careful design decisions. First, for data efficiency, the
algorithm trains a critic that is trained with dynamic programming. Now, how
can we use this critic for offline training while avoiding the bootstrapping
problem, while also avoiding modeling the data distribution, which may be
unstable? For avoiding bootstrapping error, we optimize the following problem:</p>

<p>We can compute the optimal solution for this equation and project our policy
onto it, which results in the following actor update:</p>

<p>This results in an intuitive actor update, that is also very effective in
practice. The update resembles weighted behavior cloning; if the Q function was
uninformative, it reduces to behavior cloning the replay buffer. But with a
well-formed Q estimate, we weight the policy towards only good actions. An
illustration is given in the figure above: the agent regresses onto
high-advantage actions with a large weight, while almost ignoring low-advantage
actions. Please see the paper for an expanded derivation and implementation
details.</p>

<h1>Experiments</h1>

<p>So how well does this actually do at addressing our concerns from earlier? In
our experiments, we show that we can learn difficult, high-dimensional, sparse
reward dexterous manipulation problems from human demonstrations and off-policy
data. We then evaluate our method with suboptimal prior data generated by a
random controller. Results on standard MuJoCo benchmark environments
(HalfCheetah, Walker, and Ant) are also included in the paper.</p>

<h2>Dexterous Manipulation</h2>

<p>
<img src="https://bair.berkeley.edu/static/blog/awac/14_fig5.gif" height="170"><span></span>
<img src="https://bair.berkeley.edu/static/blog/awac/15_fig5.gif" height="170"><span></span>
<img src="https://bair.berkeley.edu/static/blog/awac/16_fig5.gif" height="170"><img src="https://bair.berkeley.edu/static/blog/awac/17_fig5.png" width="100%"><br><i>Figure 5. Top: performance shown for various methods after online training
(pen: 200K steps, door: 300K steps, relocate: 5M steps). Bottom: learning
curves on dextrous manipulation tasks with sparse rewards are shown. Step 0
corresponds to the start of online training after offline pre-training.</i>
</p>

<p>We aim to study tasks representative of the difficulties of real-world robot
learning, where offline learning and online fine-tuning are most relevant. One
such setting is the suite of dexterous manipulation tasks proposed by <a href="https://arxiv.org/abs/1709.10087" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Rajeswaran et al., 2017</a>. These
tasks involve complex manipulation skills using a 28-DoF five-fingered hand in
the MuJoCo simulator: in-hand rotation of a pen, opening a door by unlatching
the handle, and picking up a sphere and relocating it to a target location.
These environments exhibit many challenges: high dimensional action spaces,
complex manipulation physics with many intermittent contacts, and randomized
hand and object positions. The reward functions in these environments are
binary 0-1 rewards for task completion. Rajeswaran et al. provide 25 human
demonstrations for each task, which are not fully optimal but do solve the
task. Since this dataset is very small, we generated another 500 trajectories
of interaction data by constructing a behavioral cloned policy, and then
sampling from this policy.</p>

<p>First, we compare our method on the dexterous manipulation tasks described
earlier against prior methods for off-policy learning, offline learning, and
bootstrapping from demonstrations. The results are shown in the figure above.
Our method uses the prior data to quickly attain good performance, and the
efficient off-policy actor-critic component of our approach fine-tunes much
quicker than DAPG. For example, our method solves the pen task in 120K
timesteps, the equivalent of just 20 minutes of online interaction. While the
baseline comparisons and ablations are able to make some amount of progress on
the pen task, alternative off-policy RL and offline RL algorithms are largely
unable to solve the door and relocate task in the time-frame considered. We
find that the design decisions to use off-policy critic estimation allow AWAC
to significantly outperform AWR while the implicit behavior modeling allows
AWAC to significantly outperform ABM, although ABM does make some progress.</p>

<h2>Fine-Tuning from Random Policy Data</h2>

<p>An advantage of using off-policy RL for reinforcement learning is that we can
also incorporate suboptimal data, rather than only demonstrations. In this
experiment, we evaluate on a simulated tabletop pushing environment with a
Sawyer robot.</p>

<p><img src="https://bair.berkeley.edu/static/blog/awac/18_random.png" height="250" hspace="40" align="right"></p>

<p>To study the potential to learn from suboptimal data, we use an off-policy
dataset of 500 trajectories generated by a random process. The task is to push
an object to a target location in a 40cm x 20cm goal space.</p>

<p>The results are shown in the figure to the right. We see that while many
methods begin at the same initial performance, AWAC learns the fastest online
and is actually able to make use of the offline dataset effectively as opposed
to some methods which are completely unable to learn.</p>

<h1>Future Directions</h1>

<p>Being able to use prior data and fine-tune quickly on new problems opens up
many new avenues of research. We are most excited about using AWAC to move from
the single-task regime in RL to the multi-task regime, with data sharing and
generalization between tasks. The strength of deep learning has been its
ability to generalize in open-world settings, which we have already seen
transform the fields of computer vision and natural language processing. To
achieve the same type of generalization in robotics, we will need RL algorithms
that take advantage of vast amounts of prior data. But one key distinction in
robotics is that collecting high-quality data for a task is very difficult -
often as difficult as solving the task itself. This is opposed to, for instance
computer vision, where humans can label the data. Thus, the active data
collection (online learning) will be an important piece of the puzzle.</p>

<p><img src="https://bair.berkeley.edu/static/blog/awac/19_future.png" height="250" hspace="40" align="right"></p>

<p>This work also suggests a number of algorithmic directions to move forward.
Note that in this work we focused on mismatched action distributions between
the policy $\pi$ and the behavior data $\pi_\beta$. When doing off-policy
learning, there is also a mismatched marginal state distribution between the
two. Intuitively, consider a problem with two solutions A and B, with B being a
higher return solution and off-policy data demonstrating solution A provided.
Even if the robot discovers solution B during online exploration, the
off-policy data still consists of mostly data from path A. Thus the Q-function
and policy updates are computed over states encountered while traversing path A
even though it will not encounter these states when executing the optimal
policy. This problem has been studied <a href="https://arxiv.org/abs/1906.04733" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">previously</a>. Accounting for both
types of distribution mismatch will likely result in better RL algorithms.</p>

<p>Finally, we are already using AWAC as a tool to speed up our research. When we
set out to solve a task, we do not usually try to solve it from scratch with
RL. First, we may teleoperate the robot to confirm the task is solvable; then
we might run some hard-coded policy or behavioral cloning experiments to see if
simple methods can already solve it. With AWAC, we can save all of the data in
these experiments, as well as other experimental data such as when
hyperparameter sweeping an RL algorithm, and use it as prior data for RL.</p>

<hr><p>A preprint of the work this blog post is based on is available <a href="https://arxiv.org/abs/2006.09359" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>. Code is now included in <a href="https://github.com/vitchyr/rlkit" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">rlkit</a>. The code documentation also
contains links to the data and environments we used. The project website is
available <a href="https://awacrl.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>.</p>]]></description>
										<content:encoded><![CDATA[<p><meta name="twitter:title" content="AWAC: Accelerating Online Reinforcement Learning with Offline Datasets" />  <meta name="twitter:card" content="summary_image" />  <meta name="twitter:image" content="https://bair.berkeley.edu/static/blog/awac/05_fig1.png" /><br />
<strong>By Ashvin Nair and Abhishek Gupta </strong></p>
<p>Robots trained with reinforcement learning (RL) have the potential to be used across a huge variety of challenging real world problems. To apply RL to a new problem, you typically set up the environment, define a reward function, and train the robot to solve the task by allowing it to explore the new environment from scratch. While this may eventually work, these “online” RL methods are data hungry and repeating this data inefficient process for every new problem makes it difficult to apply online RL to real world robotics problems. What if instead of repeating the data collection and learning process from scratch every time, we were able to reuse data across multiple problems or experiments? By doing so, we could greatly reduce the burden of data collection with every new problem that is encountered. <span id="more-194645"></span>  </p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/01_hand.gif" height="190" width="" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/02_hand.gif" height="190" width="" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/03_hand.gif" height="190" width="" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/04_hand.gif" height="190" width="" /> <br /> <br />
<i>Our method learns complex behaviors by training offline from prior datasets (expert demonstrations, data from previous experiments, or random exploration data) and then fine-tuning quickly with online interaction.</i> </p>
<p>With hundreds to thousands of robot experiments being constantly run, it is of crucial importance to devise an RL paradigm that can effectively use the large amount of already available data while still continuing to improve behavior on new tasks.</p>
<p>The first step towards moving RL towards a data driven paradigm is to consider the general idea of offline (batch) RL. Offline RL considers the problem of learning optimal policies from arbitrary off-policy data, without any further exploration. This is able to eliminate the data collection problem in RL, and incorporate data from arbitrary sources including other robots or teleoperation. However, depending on the quality of available data and the problem being tackled, we will often need to augment offline training with targeted online improvement. This problem setting actually has unique challenges of its own. In this blog post, we discuss how we can move RL from training from scratch with every new problem to a paradigm which is able to reuse prior data effectively, with some offline training followed by online finetuning.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/05_fig1.png" width="90%" /> <br /> <i>Figure 1: The problem of accelerating online RL with offline datasets. In (1), the robot learns a policy entirely from an offline dataset. In (2), the robot gets to interact with the world and collect on-policy samples to improve the policy beyond what it could learn offline.</i> </p>
<h1 id="challenges-in-offline-rl-with-online-fine-tuning">Challenges in Offline RL with Online Fine-tuning</h1>
<p>We analyze the challenges in the problem of learning from offline data and subsequent fine-tuning, using the standard benchmark HalfCheetah locomotion task. The following experiments are conducted with a prior dataset consisting of 15 demonstrations from an expert policy and 100 suboptimal trajectories sampled from a behavioral clone of these demonstrations.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/06_fig2.png" height="250" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/07_fig2.png" height="250" /> <br /> <i>Figure 2: On-policy methods are slow to learn compared to off-policy methods, due to the ability of off-policy methods to “stitch&#8221; good trajectories together, illustrated on the left. Right: in practice, we see slow online improvement using on-policy methods.</i> </p>
<h2 id="1-data-efficiency">1. Data Efficiency</h2>
<p>A simple way to utilize prior data such as demonstrations for RL is to pre-train a policy with imitation learning, and fine-tune with on-policy RL algorithms such as <a href="https://arxiv.org/abs/1910.00177" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">AWR</a> or <a href="https://arxiv.org/abs/1709.10087" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DAPG</a>. This has two drawbacks.  First, the prior data may not be optimal so imitation learning may be ineffective. Second, on-policy fine-tuning is data inefficient as it does not reuse the prior data in the RL stage. For real-world robotics, data efficiency is vital. Consider the robot on the right, trying to reach the goal state with prior trajectory $\tau_1$ and $\tau_2$. On-policy methods cannot effectively use this data, but off-policy algorithms that do dynamic programming can, by effectively “stitching” $\tau_1$ and $\tau_2$ together with the use of a value function or model. This effect can be seen in the learning curves in Figure 2, where on-policy methods are an order of magnitude slower than off-policy actor-critic methods.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/08_fig3.png" height="220" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/09_fig3.png" height="220" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/10_fig3.png" height="220" /> <br /> <i>Figure 3: Bootstrapping error is an issue when using off-policy RL for offline training. Left: an erroneous Q value far away from the data is exploited by the policy, resulting in a poor update of the Q function. Middle: as a result, the robot may take actions that are out of distribution. Right: bootstrap error causes poor offline pretraining when using SAC and its variants.</i> </p>
<h2 id="2-bootstrapping-error">2. Bootstrapping Error</h2>
<p>Actor-critic methods can in principle learn efficiently from off-policy data by estimating a value estimate $V(s)$ or action-value estimate $Q(s, a)$ of future returns by Bellman bootstrapping. However, when standard off-policy actor-critic methods are applied to our problem (we use <a href="https://arxiv.org/abs/1801.01290" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SAC</a>), they perform poorly, as shown in Figure 3: despite having a prior dataset in the replay buffer, these algorithms do not benefit significantly from offline training (as seen by the comparison between the SAC(scratch) and SACfD(prior) lines in Figure 3). Moreover, even if the policy is pre-trained by behavior cloning (“SACfD (pretrain)”) we still observe an initial decrease in performance.</p>
<p>This challenge can be attributed to off-policy bootstrapping error accumulation. During training, the Q estimates will not be fully accurate, particularly in extrapolating actions that are not present in the data. The policy update exploits overestimated Q values, making the estimated Q values worse. The issue is illustrated in the figure: incorrect Q values result in an incorrect update to the target Q values, which may result in the robot taking a poor action.</p>
<h2 id="3-non-stationary-behavior-models">3. Non-stationary Behavior Models</h2>
<p>Prior offline RL algorithms such as <a href="https://arxiv.org/pdf/1812.02900.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BCQ</a>, <a href="https://arxiv.org/abs/1906.00949" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BEAR</a>, and <a href="https://arxiv.org/abs/1911.11361" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BRAC</a> propose to address the bootstrapping issue by preventing the policy from straying too far from the data. The key idea is to prevent bootstrapping error by constraining the policy $\pi$ close to the “behavior policy” $\pi_\beta$: the actions that are present in the replay buffer. The idea is illustrated in the figure below: by sampling actions from $\pi_\beta$, you avoid exploiting incorrect Q values far away from the data distribution.</p>
<img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/11_q.png" width="1424" height="1296" class="alignnone size-full" />
<p>However, $\pi_\beta$ is typically not known, especially for offline data, and must be estimated from the data itself. Many offline RL algorithms (BEAR, BCQ, <a href="https://arxiv.org/abs/2002.08396" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ABM</a>) explicitly fit a parametric model to samples from the replay buffer for the distribution $\pi_\beta$. After forming an estimate $\hat{\pi}_\beta$, prior methods implement the policy constraint in various ways, including penalties on the policy update (BEAR, BRAC) or architecture choices for sampling actions for policy training (BCQ, ABM).</p>
<p>While offline RL algorithms with constraints perform well offline, they struggle to improve with fine-tuning, as shown in the third plot in Figure 1. We see that the purely offline RL performance (at “0K” in Fig.1) is much better than SAC. However, with additional iterations of online fine-tuning, the performance increases very slowly (as seen from the slope of the BEAR curve in Fig 1). What causes this phenomenon?</p>
<p>The issue is in fitting an accurate behavior model as data is collected online during fine-tuning. In the offline setting, behavior models must only be trained once, but in the online setting, the behavior model must be updated online to track incoming data. Training density models online (in the “streaming” setting) is a challenging research problem, made more difficult by a potentially complex multi-modal behavior distribution induced by the mixture of online and offline data. In order to address our problem setting, we require an off-policy RL algorithm that constrains the policy to prevent offline instability and error accumulation, but is not so conservative that it prevents online fine-tuning due to imperfect behavior modeling. Our proposed algorithm, which we discuss in the next section, accomplishes this by employing an implicit constraint, which does not require any explicit modeling of the behavior policy.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/12_fig4.png" height="300" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/13_fig4.png" height="300" /> <br /> <i>Figure 4: an illustration of AWAC. High-advantage transitions are regressed on with high weight, while low advantage transitions have low weight. Right: algorithm pseudocode.</i> </p>
<h1 id="advantage-weighted-actor-critic">Advantage Weighted Actor Critic</h1>
<p>In order to avoid these issues, we propose an extremely simple algorithm &#8211; advantage weighted actor critic (AWAC).  AWAC avoids the pitfalls in the previous section with careful design decisions. First, for data efficiency, the algorithm trains a critic that is trained with dynamic programming. Now, how can we use this critic for offline training while avoiding the bootstrapping problem, while also avoiding modeling the data distribution, which may be unstable? For avoiding bootstrapping error, we optimize the following problem:</p>
<p>  <script type="math/tex; mode=display">\arg\max_\pi \;  \mathbb{E}_{\mathbf{a} \sim \pi(\cdot|\mathbf{s})}[A^{\pi_k}(\mathbf{s}, \mathbf{a})] \; \text{s.t.} \; D_{\mathrm{KL}}(\pi(\cdot|\mathbf{s})||\pi_\beta(\cdot|\mathbf{s})) \leq \epsilon.</script>  </p>
<p>We can compute the optimal solution for this equation and project our policy onto it, which results in the following actor update:</p>
<p>  <script type="math/tex; mode=display">\arg\max_\theta \; \;  \mathbb{E}_{\mathbf{s}, \mathbf{a} \sim \beta}     \left[\log \pi_\theta(\mathbf{a}|\mathbf{s}) \frac{1}{Z(\mathbf{s})}  \exp \left(\frac{1}{\lambda}A^{\pi_k}(\mathbf{s}, \mathbf{a}) \right)\right].</script>  </p>
<p>This results in an intuitive actor update, that is also very effective in practice. The update resembles weighted behavior cloning; if the Q function was uninformative, it reduces to behavior cloning the replay buffer. But with a well-formed Q estimate, we weight the policy towards only good actions. An illustration is given in the figure above: the agent regresses onto high-advantage actions with a large weight, while almost ignoring low-advantage actions. Please see the paper for an expanded derivation and implementation details.</p>
<h1 id="experiments">Experiments</h1>
<p>So how well does this actually do at addressing our concerns from earlier? In our experiments, we show that we can learn difficult, high-dimensional, sparse reward dexterous manipulation problems from human demonstrations and off-policy data. We then evaluate our method with suboptimal prior data generated by a random controller. Results on standard MuJoCo benchmark environments (HalfCheetah, Walker, and Ant) are also included in the paper.</p>
<h2 id="dexterous-manipulation">Dexterous Manipulation</h2>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/14_fig5.gif" height="170" /> <span class="vertical-line-170p"></span> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/15_fig5.gif" height="170" /> <span class="vertical-line-170p"></span> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/16_fig5.gif" height="170" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/17_fig5.png" width="100%" /> <br /> <i>Figure 5. Top: performance shown for various methods after online training (pen: 200K steps, door: 300K steps, relocate: 5M steps). Bottom: learning curves on dextrous manipulation tasks with sparse rewards are shown. Step 0 corresponds to the start of online training after offline pre-training.</i> </p>
<p>We aim to study tasks representative of the difficulties of real-world robot learning, where offline learning and online fine-tuning are most relevant. One such setting is the suite of dexterous manipulation tasks proposed by <a href="https://arxiv.org/abs/1709.10087" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Rajeswaran et al., 2017</a>. These tasks involve complex manipulation skills using a 28-DoF five-fingered hand in the MuJoCo simulator: in-hand rotation of a pen, opening a door by unlatching the handle, and picking up a sphere and relocating it to a target location. These environments exhibit many challenges: high dimensional action spaces, complex manipulation physics with many intermittent contacts, and randomized hand and object positions. The reward functions in these environments are binary 0-1 rewards for task completion. Rajeswaran et al. provide 25 human demonstrations for each task, which are not fully optimal but do solve the task. Since this dataset is very small, we generated another 500 trajectories of interaction data by constructing a behavioral cloned policy, and then sampling from this policy.</p>
<p>First, we compare our method on the dexterous manipulation tasks described earlier against prior methods for off-policy learning, offline learning, and bootstrapping from demonstrations. The results are shown in the figure above. Our method uses the prior data to quickly attain good performance, and the efficient off-policy actor-critic component of our approach fine-tunes much quicker than DAPG. For example, our method solves the pen task in 120K timesteps, the equivalent of just 20 minutes of online interaction. While the baseline comparisons and ablations are able to make some amount of progress on the pen task, alternative off-policy RL and offline RL algorithms are largely unable to solve the door and relocate task in the time-frame considered. We find that the design decisions to use off-policy critic estimation allow AWAC to significantly outperform AWR while the implicit behavior modeling allows AWAC to significantly outperform ABM, although ABM does make some progress.</p>
<h2 id="fine-tuning-from-random-policy-data">Fine-Tuning from Random Policy Data</h2>
<p>An advantage of using off-policy RL for reinforcement learning is that we can also incorporate suboptimal data, rather than only demonstrations. In this experiment, we evaluate on a simulated tabletop pushing environment with a Sawyer robot.</p>
<img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/18_random.png" height="250" hspace="40" align="right" />
<p>To study the potential to learn from suboptimal data, we use an off-policy dataset of 500 trajectories generated by a random process. The task is to push an object to a target location in a 40cm x 20cm goal space.</p>
<p>The results are shown in the figure to the right. We see that while many methods begin at the same initial performance, AWAC learns the fastest online and is actually able to make use of the offline dataset effectively as opposed to some methods which are completely unable to learn.</p>
<h1 id="future-directions">Future Directions</h1>
<p>Being able to use prior data and fine-tune quickly on new problems opens up many new avenues of research. We are most excited about using AWAC to move from the single-task regime in RL to the multi-task regime, with data sharing and generalization between tasks. The strength of deep learning has been its ability to generalize in open-world settings, which we have already seen transform the fields of computer vision and natural language processing. To achieve the same type of generalization in robotics, we will need RL algorithms that take advantage of vast amounts of prior data. But one key distinction in robotics is that collecting high-quality data for a task is very difficult &#8211; often as difficult as solving the task itself. This is opposed to, for instance computer vision, where humans can label the data. Thus, the active data collection (online learning) will be an important piece of the puzzle.</p>
<img decoding="async" src="https://bair.berkeley.edu/static/blog/awac/19_future.png" height="250" hspace="40" align="right" />
<p>This work also suggests a number of algorithmic directions to move forward. Note that in this work we focused on mismatched action distributions between the policy $\pi$ and the behavior data $\pi_\beta$. When doing off-policy learning, there is also a mismatched marginal state distribution between the two. Intuitively, consider a problem with two solutions A and B, with B being a higher return solution and off-policy data demonstrating solution A provided. Even if the robot discovers solution B during online exploration, the off-policy data still consists of mostly data from path A. Thus the Q-function and policy updates are computed over states encountered while traversing path A even though it will not encounter these states when executing the optimal policy. This problem has been studied <a href="https://arxiv.org/abs/1906.04733" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">previously</a>. Accounting for both types of distribution mismatch will likely result in better RL algorithms.</p>
<p>Finally, we are already using AWAC as a tool to speed up our research. When we set out to solve a task, we do not usually try to solve it from scratch with RL. First, we may teleoperate the robot to confirm the task is solvable; then we might run some hard-coded policy or behavioral cloning experiments to see if simple methods can already solve it. With AWAC, we can save all of the data in these experiments, as well as other experimental data such as when hyperparameter sweeping an RL algorithm, and use it as prior data for RL.</p>
<hr />
<p>A preprint of the work this blog post is based on is available <a href="https://arxiv.org/abs/2006.09359" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>. Code is now included in <a href="https://github.com/vitchyr/rlkit" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">rlkit</a>. The code documentation also contains links to the data and environments we used. The project website is available <a href="https://awacrl.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>.</p>
<p>This article was initially published on the <a href="http://bair.berkeley.edu/blog/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Can RL from pixels be as efficient as RL from state?</title>
		<link>https://robohub.org/can-rl-from-pixels-be-as-efficient-as-rl-from-state/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 16 Sep 2020 20:22:27 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/can-rl-from-pixels-be-as-efficient-as-rl-from-state/</guid>

					<description><![CDATA[By Misha Laskin, Aravind Srinivas, Kimin Lee, Adam Stooke, Lerrel Pinto, Pieter Abbeel A remarkable characteristic of human intelligence is our ability to learn tasks quickly. Most humans can learn reasonably complex skills like tool-use and gameplay within just a few hours, and understand the basics after only a few attempts. This suggests that data-efficient [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" src="https://robohub.org/wp-content/uploads/2020/09/Combining-contrastive-learning.png" alt="" width="2400" height="1133" class="aligncenter size-full wp-image-195007" srcset="https://robohub.org/wp-content/uploads/2020/09/Combining-contrastive-learning.png 2400w, https://robohub.org/wp-content/uploads/2020/09/Combining-contrastive-learning-425x201.png 425w, https://robohub.org/wp-content/uploads/2020/09/Combining-contrastive-learning-1024x483.png 1024w, https://robohub.org/wp-content/uploads/2020/09/Combining-contrastive-learning-768x363.png 768w, https://robohub.org/wp-content/uploads/2020/09/Combining-contrastive-learning-1536x725.png 1536w, https://robohub.org/wp-content/uploads/2020/09/Combining-contrastive-learning-2048x967.png 2048w" sizes="(max-width: 2400px) 100vw, 2400px" /><br />
By <strong><a href="https://mishalaskin.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Misha Laskin</a>, <a href="https://people.eecs.berkeley.edu/~aravind/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Aravind Srinivas</a>, <a href="https://sites.google.com/view/kiminlee" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Kimin Lee</a>, <a href="https://www.linkedin.com/in/adam-stooke-06bb6923/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Adam Stooke</a>, <a href="https://cs.nyu.edu/~lp91/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Lerrel Pinto</a>, <a href="https://people.eecs.berkeley.edu/~pabbeel/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Pieter Abbeel</a></strong></p>
<p>A remarkable characteristic of human intelligence is our ability to learn tasks quickly. Most humans can learn reasonably complex skills like tool-use and gameplay within just a few hours, and understand the basics after only a few attempts. This suggests that data-efficient learning may be a meaningful part of developing broader intelligence.</p>
<p><span id="more-195004"></span></p>
<p>On the other hand, Deep Reinforcement Learning (RL) algorithms can achieve superhuman performance on games like Atari, Starcraft, Dota, and Go, but require large amounts of data to get there. Achieving superhuman performance on Dota took over <em>10,000 human years</em> of gameplay. Unlike simulation, skill acquisition in the real-world is constrained to wall-clock time. In order to see similar breakthroughs to AlphaGo in real-world settings, such as robotic manipulation and autonomous vehicle navigation, RL algorithms need to be data-efficient — they need to learn effective policies within a reasonable amount of time.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/curl/fig1.png" width="80%" />
</p>
<p>To date, it has been commonly assumed that RL operating on coordinate state is significantly more data-efficient than pixel-based RL. However, coordinate state is just a human crafted representation of visual information. In principle, if the environment is fully observable, we should also be able to learn representations that capture the state.</p>
<h1 id="recent-advances-in-data-efficient-rl">Recent advances in data-efficient RL</h1>
<p>Recently, there have been several algorithmic advances in Deep RL that have improved learning policies from pixels. The methods fall into two categories: (i) model-free algorithms and (ii) model-based (MBRL) algorithms. The main difference between the two is that model-based methods learn a forward transition model $p(s_{t+1}|,s_t,a_t)$ while model-free ones do not. Learning a model has several distinct advantages. First, it is possible to use the model to plan through action sequences, generate fictitious rollouts as a form of data augmentation, and temporally shape the latent space by learning a model.</p>
<p>However, a distinct disadvantage of model-based RL is complexity. Model-based methods operating on pixels require learning a model, an encoding scheme, a policy, various auxiliary tasks such as reward prediction, and stitching these parts together to make a whole algorithm. Visual MBRL methods have a lot of moving parts and tend to be less stable. On the other hand, model-free methods such as Deep Q Networks (DQN), Proximal Policy Optimization (PPO), and Soft Actor-Critic (SAC) learn a policy in an end-to-end manner optimizing for one objective. While traditionally, the simplicity of model-free RL has come at the cost of sample-efficiency, recent improvements have shown that model-free methods can in fact be more data-efficient than MBRL and, more surprisingly, result in policies that are as data efficient as policies trained on coordinate state. In what follows we will focus on these recent advances in pixel-based model-free RL.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/curl/fig2.png" width="" />
</p>
<h1 id="why-now">Why now?</h1>
<p>Over the last few years, two trends have converged to make data-efficient visual RL possible. First, end-to-end RL algorithms have become increasingly more stable through algorithms like the Rainbow DQN, TD3, and SAC. Second, there has been tremendous progress in label-efficient learning for image classification using contrastive unsupervised representations (CPCv2, MoCo, SimCLR) and data augmentation (MixUp, AutoAugment, RandAugment). In recent work from our lab at BAIR (CURL, RAD), we combined contrastive learning and data augmentation techniques from computer vision with model-free RL to show significant data-efficiency gains on common RL benchmarks like Atari, DeepMind control, ProcGen, and OpenAI gym.</p>
<h1 id="contrastive-learning-in-rl-setting">Contrastive Learning in RL Setting</h1>
<p>CURL was inspired by recent advances in contrastive representation learning in computer vision (CPC, CPCv2, MoCo, SimCLR). Contrastive learning aims to maximize / minimize similarity between two similar / dissimilar representations of an image. For example, in MoCo and SimCLR, the objective is to maximize agreement between two data-augmented versions of the same image and minimize it between all other images in the dataset, where optimization is performed with a Noise Contrastive Estimation loss. Through data augmentation, these representations internalize powerful inductive biases about invariance in the dataset.</p>
<p>In the RL setting, we opted for a similar approach and adopted the momentum contrast (MoCo) mechanism, a popular contrastive learning method in computer vision that uses a moving average of the query encoder parameters (momentum) to encode the keys to stabilize training. There are two main differences in setup: (i) the RL dataset changes dynamically and (ii) visual RL is typically performed on stacks of frames to access temporal information like velocities. Rather than separating contrastive learning from the downstream task as done in vision, we learn contrastive representations jointly with the RL objective. Instead of discriminating across single images, we discriminate across the stack of frames.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/curl/fig3.png" width="" />
</p>
<p>By combining contrastive learning with Deep RL in the above manner <em>we found, for the first time, that pixel-based RL can be nearly as data-efficient as state-based RL</em> on the DeepMind control benchmark suite. In the figure below, we show learning curves for DeepMind control tasks where contrastive learning is coupled with SAC (red) and compared to state-based SAC (gray).</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/curl/fig4.png" width="" />
</p>
<p>We also demonstrate data-efficiency gains on the Atari 100k step benchmark. In this setting, we couple CURL with an Efficient Rainbow DQN (Eff. Rainbow) and show that CURL outperforms the prior state-of-the-art (Eff. Rainbow, SimPLe) on 20 out of 26 games tested.</p>
<h1 id="rl-with-data-augmentation">RL with Data Augmentation</h1>
<p>Given that random cropping was a crucial component in CURL, it is natural to ask — can we achieve the same results with data augmentation alone? In Reinforcement Learning with Augmented Data (RAD), we performed the first extensive study of data augmentation in Deep RL and found that for the DeepMind control benchmark the answer is yes. Data augmentation alone can outperform prior competing methods, match, and sometimes surpass the efficiency of state-based RL. Similar results were also shown in concurrent work &#8211; DrQ.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/curl/fig5.png" width="" />
</p>
<p>We found that RAD also improves generalization on the ProcGen game suite, showing that data augmentation is not limited to improving data-efficiency but also helps RL methods generalize to test-time environments.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/curl/fig6.png" width="" />
</p>
<p>If data augmentation works for pixel-based RL, can it also improve state-based methods? We introduced a new state-based augmentation — <em>random amplitude scaling</em> — and showed that simple RL with state-based data augmentation achieves state-of-the-art results on OpenAI gym environments and outperforms more complex model-free and model-based RL algorithms.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/curl/fig7.png" width="" />
</p>
<h1 id="contrastive-learning-vs-data-augmentation">Contrastive Learning vs Data Augmentation</h1>
<p>If data augmentation with RL performs so well, do we need unsupervised representation learning? RAD outperforms CURL because it only optimizes for what we care about, which is the task reward. CURL, on the other hand, jointly optimizes the reinforcement and contrastive learning objectives. If the metric<br />
used to evaluate and compare these methods is the score attained on the task at hand, a method that purely focuses on reward optimization is expected to be better as long as it implicitly ensures similarity consistencies on the augmented views.</p>
<p>However, many problems in RL cannot be solved with data augmentations alone. For example, RAD would not be applicable to environments with sparse-rewards or no rewards at all, because it learns similarity consistency implicitly through the observations coupled to a reward signal. On the other hand, the contrastive learning objective in CURL internalizes invariances explicitly and is therefore able to learn semantic representations from high dimensional observations gathered from any rollout regardless of the reward signal. Unsupervised representation learning may therefore be a better fit for real-world tasks, such as robotic manipulation, where the environment reward is more likely to be sparse or absent.</p>
<hr />
<p>This post is based on the following papers:</p>
<ul>
<li>
<p><strong>CURL: Contrastive Unsupervised Representations for Reinforcement Learning</strong><br />
Michael Laskin*, Aravind Srinivas*, Pieter Abbeel<br />
Thirty-seventh International Conference Machine Learning (ICML), 2020.<br />
<a href="https://arxiv.org/abs/2004.04136" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">arXiv</a>, <a href="https://mishalaskin.github.io/curl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project Website</a></p>
</li>
<li>
<p><strong>Reinforcement Learning with Augmented Data</strong><br />
Michael Laskin*, Kimin Lee*, Adam Stooke, Lerrel Pinto, Pieter Abbeel, Aravind Srinivas<br />
<a href="https://arxiv.org/abs/2004.14990" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">arXiv</a>, <a href="https://mishalaskin.github.io/rad/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project Website</a></p>
</li>
</ul>
<p><strong>References</strong></p>
<ol><font size="-1"></p>
<li>Hafner et al. <a href="https://arxiv.org/abs/1811.04551" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Learning Latent Dynamics for Planning from Pixels</a>. ICML 2019.</li>
<li>Hafner et al. <a href="https://arxiv.org/abs/1912.01603" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Dream to Control: Learning Behaviors by Latent Imagination</a>. ICLR 2020.</li>
<li>Kaiser et al. <a href="https://arxiv.org/abs/1903.00374" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Model-Based Reinforcement Learning for Atari</a>. ICLR 2020.</li>
<li>Lee et al. <a href="https://arxiv.org/abs/1907.00953" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable Model</a>. arXiv 2019.</li>
<li>Henaff et al. <a href="https://arxiv.org/abs/1905.09272" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Data-Efficient Image Recognition with Contrastive Predictive Coding</a>. ICML 2020.</li>
<li>He et al. <a href="https://arxiv.org/abs/1911.05722" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Momentum Contrast for Unsupervised Visual Representation Learning</a>. CVPR 2020.</li>
<li>Chen et al. <a href="https://arxiv.org/abs/2002.05709" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">A Simple Framework for Contrastive Learning of Visual Representations</a>. ICML 2020.</li>
<li>Kostrikov et al. <a href="https://arxiv.org/abs/2004.13649" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels</a>. arXiv 2020.</li>
<p></font></ol>
<p>This article was initially published on the <a href="https://bair.berkeley.edu/blog/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>OmniTact: a multi-directional high-resolution touch sensor</title>
		<link>https://robohub.org/omnitact-a-multi-directional-high-resolution-touch-sensor/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 22 Jul 2020 21:36:22 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/omnitact-a-multi-directional-high-resolution-touch-sensor/</guid>

					<description><![CDATA[By Akhil Padmanabha and Frederik Ebert Touch has been shown to be important for dexterous manipulation in robotics. Recently, the GelSight sensor has caught significant interest for learning-based robotics due to its low cost and rich signal. For example, GelSight sensors have been used for learning inserting USB cables (Li et al, 2014), rolling a [&#8230;]]]></description>
										<content:encoded><![CDATA[<p><div id="attachment_179352" style="width: 910px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-179352" src="https://robohub.org/wp-content/uploads/2020/07/OmniTact-image.png" alt="" width="900" height="436" class="size-full wp-image-179352" srcset="https://robohub.org/wp-content/uploads/2020/07/OmniTact-image.png 900w, https://robohub.org/wp-content/uploads/2020/07/OmniTact-image-425x206.png 425w, https://robohub.org/wp-content/uploads/2020/07/OmniTact-image-768x372.png 768w" sizes="(max-width: 900px) 100vw, 900px" /><p id="caption-attachment-179352" class="wp-caption-text">Human thumb next to our OmniTact sensor, and a US penny for scale.</p></div><br />
<strong>By Akhil Padmanabha and Frederik Ebert</strong></p>
<p><a href="https://bair.berkeley.edu/blog/2019/03/21/tactile/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Touch has been shown</a> to be important for dexterous <a href="https://www.researchgate.net/profile/Ravinder_Dahiya2/publication/221787495_Tactile_Sensing_for_ obotic_Applications/links/0fcfd50ff8e9185522000000.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">manipulation</a> in <a href="http://www.biorobotics.harvard.edu/pubs/1994/tac-manip.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robotics</a>. Recently, the GelSight sensor has caught significant interest for <em>learning-based robotics</em> due to its low cost and rich signal. For example, GelSight sensors have been used for learning inserting USB cables (<a href="https://dspace.mit.edu/bitstream/handle/1721.1/88136/GelSight_IROS%202014_final.pdf sequence=1&amp;isAllowed=y" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Li et al, 2014</a>), rolling a die (<a href="https://arxiv.org/abs/1903.04128" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Tian et al. 2019</a>) or grasping objects (<a href="https://arxiv.org/abs/1710.05512" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Calandra et al.  2017</a>).</p>
<p><span id="more-179351"></span></p>
<p>The reason why learning-based methods work well with GelSight sensors is that they output high-resolution tactile images from which a variety of features such as <a href="https://www.gelsight.com/applications/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">object geometry</a>, surface texture, normal and shear forces can be <a href="http://people.csail.mit.edu/yuan_wz/force-torque.htm" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">estimated</a> that often prove critical to robotic control. The tactile images can be fed into standard CNN-based computer vision pipelines allowing the use of a variety of different learning-based techniques: In <a href="https://arxiv.org/abs/1710.05512" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Calandra et al. 2017</a> a grasp-success classifier is trained on GelSight data collected in self-supervised manner, in <a href="https://arxiv.org/abs/1903.04128" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Tian et al. 2019</a> <a href="https://bair.berkeley.edu/blog/2018/11/30/visual-rl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Visual Foresight</a>, a video-prediction-based control algorithm is used to make a robot roll a die purely based on tactile images, and in <a href="https://ieeexplore.ieee.org/abstract/document/9018215" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Lambeta et al. 2020</a> a model-based RL algorithm is applied to in-hand manipulation using GelSight images.</p>
<p>Unfortunately applying GelSight sensors in practical real-world scenarios is still challenging due to its large size and the fact that it is only sensitive on one side. Here we introduce a new, more compact tactile sensor design based on GelSight that allows for omnidirectional sensing, i.e. making the sensor <em>sensitive on all sides like a human finger</em>, and show how this opens up new possibilities for sensorimotor learning. We demonstrate this by teaching a robot to pick up electrical plugs and insert them <em>purely based on tactile feedback</em>.</p>
<h1 id="gelsight-sensors">GelSight Sensors</h1>
<p>A standard GelSight sensor, shown in the figure below on the left, uses an off-the-shelf webcam to capture high-resolution images of deformations on the silicone gel skin. The inside surface of the gel skin is illuminated with colored LEDs, providing sufficient lighting for the tactile image.</p>
<p style="text-align:center;">
<figure><img decoding="async" src="https://bair.berkeley.edu/static/blog/omnitact/Fig2.svg" width="500" /><figcaption><i><br />
Comparison of GelSight-style sensor (left side) to our OmniTact sensor (right side).<br />
</i></figcaption></figure>
</p>
<p>Existing GelSight designs are either flat, have small sensitive fields or only provide low-resolution signals. For example, prior versions of the GelSight sensor, provide high resolution (400&#215;400 pixel) images but are large and flat, providing sensitivity on only one side, while the commercial <a href="https://onrobot.com/en/products/hex-6-axis-force-torque-sensor" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">OptoForce</a> sensor (recently discontinued by OnRobot) is curved, but only provides force readings as a single 3-dimensional force vector.</p>
<h1 id="the-omnitact-sensor">The OmniTact Sensor</h1>
<p>Our OmniTact sensor design aims to address these limitations. It provides both multi-directional and high-resolution sensing on its curved surface in a compact form factor.  Similar to GelSight, OmniTact uses cameras embedded into a silicone gel skin to capture deformation of the skin, providing a rich signal from which a wide range of features such as shear and normal forces, object pose, geometry and material properties can be inferred. OmniTact uses multiple cameras giving it both high-resolution and multi-directional capabilities. The sensor itself can be used as a “finger” and can be integrated into a gripper or robotic hand. It is more compact than previous GelSight sensors, which is accomplished by utilizing micro-cameras typically used in endoscopes, and by casting the silicone gel directly onto the cameras. Tactile images from OmniTact are shown in the figures below.</p>
<p style="text-align:center;">
<figure><img decoding="async" src="https://bair.berkeley.edu/static/blog/omnitact/Fig3.svg" width="" /><figcaption><i><br />
Tactile readings from OmniTact with various objects. From left to right: M3 Screw Head, M3 Screw Threads, Combination Lock with numbers 4 3 9, Printed Circuit Board (PCB), Wireless Mouse USB. All images are taken from the upward-facing camera.<br />
</i></figcaption></figure>
</p>
<p style="text-align:center;">
<figure><img decoding="async" src="https://bair.berkeley.edu/static/blog/omnitact/Fig4.svg" width="" /><figcaption><i><br />
Tactile readings from the OmniTact being rolled over a gear rack. The multi-directional capabilities of OmniTact keep the gear rack in view as the sensor is rotated.<br />
</i></figcaption></figure>
</p>
<h1 id="design-highlights">Design Highlights</h1>
<p>One of our primary goals throughout the design process was to make OmniTact as compact as possible. To accomplish this goal, we used micro-cameras with large viewing angles and a small focus distance. Specifically we picked cameras that are commonly used in medical endoscopes measuring just (1.35 x 1.35 x 5 mm) in size with a focus distance of 5 mm. These cameras were arranged in a 3D printed camera mount as shown in the figure below which allowed us to minimize blind spots on the surface of the sensor and reduce the diameter (D) of the sensor to 30 mm.</p>
<p style="text-align:center;">
<figure><img decoding="async" src="https://bair.berkeley.edu/static/blog/omnitact/Fig5.svg" width="400" /><figcaption><i><br />
This image shows the fields of view and arrangement of the 5 micro-cameras inside the sensor. Using this arrangement, most of the fingertip can be made sensitive effectively. In the vertical plane, shown in A, we obtain $\alpha=270$ degrees of sensitivity. In the horizontal plane, shown in B, we obtain 360 degrees of sensitivity, except for small blind spots between the fields of view.<br />
</i></figcaption></figure>
</p>
<h1 id="electrical-connector-insertion-task">Electrical Connector Insertion Task</h1>
<p>We show that OmniTact’s multi-directional tactile sensing capabilities can be leveraged to solve a challenging robotic control problem: Inserting an electrical connector blindly into a wall outlet purely based on information from the multi-directional touch sensor (shown in the figure below). This task is challenging since it requires localizing the electrical connector relative to the gripper and localizing the gripper relative to the wall outlet.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/omnitact/Fig6.svg" width="" /><br />

</p>
<p>To learn the insertion task, we used a simple imitation learning algorithm that estimates the end-effector displacement required for inserting the plug into the outlet based on the tactile images from the OmniTact sensor. Our model was trained with just 100 demonstrations of insertion by controlling the robot using keyboard control. Successful insertions obtained by running the trained policy are shown in the gifs below.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/omnitact/gif1.gif" height="150" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/omnitact/gif2.gif" height="150" /><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/omnitact/gif3.gif" height="150" /><br />

</p>
<p>As shown in the table below, using the multi-directional capabilities (both the top and side camera) of our sensor allowed for the highest success rate (80%) in comparison to using just one camera from the sensor, indicating that multi-directional touch sensing is indeed crucial for solving this task. We additionally compared performance with another multi-directional tactile sensor, the OptoForce sensor, which only had a success rate of 17%.</p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/omnitact/Fig7.png" width="" /><br />

</p>
<h1 id="whats-next">What’s Next?</h1>
<p>We believe that compact, high resolution and multi-directional touch sensing has the potential to transform the capabilities of current robotic manipulation systems. We suspect that multi-directional tactile sensing could be an essential element in general-purpose robotic manipulation in addition to applications such as robotic teleoperation in surgery, as well as in sea and space missions. In the future, we plan to make OmniTact cheaper and more compact, allowing it to be used in a wider range of tasks. Our team additionally plans to conduct more robotic manipulation research that will inform future generations of tactile sensors.</p>
<p>This blog post is based on the following paper which will be presented at the International Conference on Robotics and Automation 2020:</p>
<ul>
<li><strong>OmniTact: A Multi-Directional High-Resolution Touch Sensor</strong><br />
Akhil Padmanabha, Frederik Ebert, Stephen Tian, Roberto Calandra, Chelsea Finn, Sergey Levine<br />
Paper Link: <a href="https://arxiv.org/abs/2003.06965" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">https://arxiv.org/abs/2003.06965</a><br />
Research Website: <a href="https://sites.google.com/berkeley.edu/omnitact/home" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">https://sites.google.com/berkeley.edu/omnitact/home</a></li>
</ul>
<p>We would like to thank Professor Sergey Levine, Professor Chelsea Finn, and Stephen Tian for their valuable feedback when preparing this blog post.</p>
<p>This article was initially published on the <a href="https://bair.berkeley.edu/blog/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Unsupervised meta-learning: learning to learn without supervision</title>
		<link>https://robohub.org/unsupervised-meta-learning-learning-to-learn-without-supervision/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 06 May 2020 21:09:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/unsupervised-meta-learning-learning-to-learn-without-supervision/</guid>

					<description><![CDATA[<!--
TODO TODO TODO personal reminder for Daniel Seita :-)
Be careful that these three lines are at the top,
and that the title and image change for each blog post!
-->
<p><em>This post is cross-listed <a href="https://blog.ml.cmu.edu/2020/05/01/unsupervised-meta-learning-learning-to-learn-without-supervision/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">on the CMU ML blog</a>.</em></p>

<p>The history of machine learning has largely been a story of increasing
abstraction. In the dawn of ML, researchers spent considerable effort
engineering features. As deep learning gained popularity, researchers then
shifted towards tuning the update rules and learning rates for their
optimizers. Recent research in meta-learning has climbed one level of
abstraction higher: many researchers now spend their days manually constructing
task distributions, from which they can automatically learn good optimizers.
What might be the next rung on this ladder? In this post we introduce theory
and algorithms for <strong>unsupervised meta-learning</strong>, where machine learning
algorithms themselves propose their own task distributions. Unsupervised
meta-learning further reduces the amount of human supervision required to solve
tasks, potentially inserting a new rung on this ladder of abstraction.</p>

<!--more-->

<p>We start by discussing how machine learning algorithms use human supervision to
find patterns and extract knowledge from observed data. The most common machine
learning setting is <em>regression</em>, where a human provides labels $Y$ for a set of
examples $X$. The aim is to return a predictor that correctly assigns
labels to novel examples. Another common machine learning problem setting is
<em>reinforcement learning (RL)</em>, where an agent takes actions in an environment.
In RL, humans indicate the desired behavior through a reward function that
the agent seeks to maximize. To draw a crude analogy to regression,
the environment dynamics are the examples $X$, and the reward function gives
the labels $Y$. Algorithms for regression and RL employ many tools,
including tabular methods (e.g., value iteration), linear methods (e.g., linear
regression) kernel-methods (e.g., RBF-SVMs), and deep neural networks. Broadly, we call these
algorithms <em>learning procedures</em>: processes that take as input a dataset
(examples with labels, or transitions with rewards) and output a function that
performs well (achieves high accuracy or large reward) on the dataset.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/umrl/control_room.jpg" width="500"><br><i>
Machine learning research is similar to the control room for large physics
experiments. Researchers have a number of knobs they can tune which affect the
performance of the learning procedure. The right setting for the knobs depends
on the particular experiment: some settings work well for high-energy
experiments; others work well for ultracold atom experiments.
<a href="https://www.flickr.com/photos/x-ray_delta_one/3941701730" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Figure Credit</a>.
</i>
</p>

<p>Similar to lab procedures used in physics and biology, the learning
procedures used in machine learning have many knobs<sup><a href="http://bair.berkeley.edu/blog/#fn:knob" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">1</a></sup> that can be tuned.
For example, the learning procedure for training a neural network might be defined by an optimizer
(e.g., Nesterov, Adam) and a learning rate (e.g., 1e-5).
Compared with regression, learning procedures specific to RL (e.g., DDPG) often have many more knobs, including
the frequency of data collection and how frequently the policy is updated.
Finding the right setting for the knobs can have a large effect on how quickly
the learning procedure solves a task, and a good configuration of knobs for one
learning procedure may be a bad configuration for another.</p>

<h1>Meta-Learning Optimizes Knobs of the Learning Procedure</h1>

<p>While machine learning practitioners often carefully tune these knobs by hand,
if we are going to solve many tasks, it may be useful to automatic this
process. The process of setting the knobs of learning procedures via
optimization is called <em>meta-learning</em> [<a href="https://link.springer.com/chapter/10.1007/978-1-4615-5529-2_1" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Thrun 1998</a>]. Algorithms that perform this optimization
problem automatically are known as <em>meta-learning algorithms</em>.
Explicitly tuning the knobs of learning procedures is an active area of research, with various researchers looking at tuning the update rules [<a href="https://arxiv.org/abs/1606.04474" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Andrychowicz 2016</a>, <a href="https://arxiv.org/abs/1611.02779" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Duan 2016</a>, <a href="https://arxiv.org/abs/1611.05763" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wang 2016</a>], weight initialization [<a href="https://arxiv.org/abs/1703.03400" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Finn 2017</a>], network weights [<a href="https://arxiv.org/abs/1609.09106" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Ha 2016</a>], network architectures [<a href="https://arxiv.org/abs/1906.04358" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gaier 2019</a>], and other facets of learning procedures.</p>

<p>To evaluate a setting of knobs, meta-learning algorithms consider not one task
but a distribution over many tasks. For example, a distribution over supervised
learning tasks may include learning a dog detector, learning a cat detector,
and learning a bird detector. In reinforcement learning, a task distribution
could be defined as driving a car in a smooth, safe, and efficient manner,
where tasks differ by the weights they place on smoothness, safety, and
efficiency. Ideally, the task distribution is designed to mirror the
distribution over tasks that we are likely to encounter in the real world.
Since the tasks in a task distribution are typically related, information from
one task may be useful in solving other tasks more efficiently. As you might
expect, a knob setting that works best on one distribution of tasks may not be
the best for another task distribution; the optimal knob setting
depends on the task distribution.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/umrl/Meta-Learning-v2.png" width=""><br><i>
An illustration of meta-learning, where tasks correspond to arranging blocks
into different types of towers. The human has a particular block tower in mind
and rewards the robot when it builds the correct tower. The robot's aim is to
build the correct tower as quickly as possible.
</i>
</p>

<p>In many settings we want to do well on a task distribution to which we have
only limited access. For example, in a self-driving car, tasks may correspond
to finding the optimal balance of smoothness, safety, and efficiency for each
rider, but querying riders to get rewards is expensive. A researcher can
attempt to manually construct a task distribution that mimics the true task
distribution, but this can be quite challenging and time consuming. Can we
avoid having to manually design such task distributions?</p>

<p>To answer this question, we must understand where the benefits of meta-learning
come from. When we define
task distributions for meta-learning, we do so with some prior knowledge in
mind. Without this prior information, tuning the knobs of a learning procedure
is often a zero-sum game: setting the knobs to any configuration will
accelerate learning on some tasks while slowing learning on other tasks. Does
this suggest there is no way to see the benefit of meta-learning without the
manual construction of task distributions? Perhaps not! The next section
presents an alternative.</p>

<h1>Optimizing the Learning Procedure with Self-Proposed Tasks</h1>

<p>If designing task distributions is the bottleneck in applying meta-learning
algorithms, why not have meta-learning algorithms propose their own tasks? At
first glance this seems like a terrible idea, because the No Free Lunch Theorem
suggests that this is impossible, <em>without additional knowledge</em>. However, many
real-world settings do provide a bit of additional information, albeit
disguised as unlabeled data. For example, in regression, we might have
access to an unlabeled dataset and know that the downstream tasks will be
labeled versions of this same image dataset. In a RL setting, a robot can
interact with its environment without receiving any reward, knowing that
downstream tasks will be constructed by defining reward functions for this very
environment (i.e. the real world). Seen from this perspective, the recipe for
<em>unsupervised meta-learning</em> (doing meta-learning without manually constructed
tasks) becomes clear: given unlabeled data, construct task distributions from this unlabeled data or environment, and then meta-learn to quickly solve these self-proposed tasks.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/umrl/Unsupervised-Meta-Learning-v2.png" width=""><br><i>
In unsupervised meta-learning, the agent proposes its own tasks, rather than
relying on tasks proposed by a human.
</i>
</p>

<p>How can we use this unlabeled data to construct task distributions which will
facilitate learning downstream tasks? In the case of regression, prior work on
unsupervised meta-learning [<a href="https://arxiv.org/abs/1810.02334" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Hsu 2018</a>, <a href="https://arxiv.org/abs/1811.11819" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Khodadadeh 2019</a>]
clusters an unlabeled dataset of images and then randomly chooses subsets of
the clusters to define a distribution of classification tasks. Other work
[<a href="https://arxiv.org/abs/1912.04226" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Jabri 2019</a>] look at an RL setting: after exploring an environment without a
reward function to collect a set of behaviors that are feasible in this
environment, these behaviors are clustered and used to define a distribution of
reward functions. In both cases, even though the tasks constructed can be
random, the resulting task distribution is not random, because all tasks share
the underlying unlabeled data &#8212; the image dataset for regression and
the environment dynamics for reinforcement learning. <em>The underlying unlabeled
data are the inductive bias with which we pay for our free lunch.</em></p>

<p>Let us take a deeper look into the RL case. Without knowing the downstream tasks or reward functions, what is the &#8220;best&#8221; task distribution for &#8220;practicing&#8221; to solve tasks quickly? Can we measure how effective a task distribution is for solving unknown,
downstream tasks? Is there any sense in which one unsupervised task
proposal mechanism is better than another? Understanding the answers to these
questions may guide the principled development of meta-learning algorithms with
little dependence on human supervision. Our work [<a href="https://arxiv.org/abs/1806.04640" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gupta 2018</a>], takes a
first step towards answering these questions. In particular, we examine the
<em>worst-case</em> performance of learning procedures, and derive an optimal
unsupervised meta-reinforcement learning procedure.</p>

<h1>Optimal Unsupervised Meta-Learning</h1>

<p>To answer the questions posed above, our first step is to define an optimal
meta-learner for the case where the distribution of tasks is known.
We define an optimal meta-learner as the learning procedure that achieves the
largest expected reward, averaged across the distribution of tasks. More precisely,
we will compare the expected reward for a learning procedure $f$
to that of best learning procedure $f^*$,
defining the <em>regret</em> of $f$ on a task distribution $p$ as follows:</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/umrl/labeled_regret_v2.png" width="600"><br></p>

<p>Extending this definition to the case of unsupervised meta-learning, an optimal
unsupervised meta-learner can be defined as a meta-learner that achieves the
minimum <em>worst-case</em> regret across all possible task distributions that may be
encountered in the environment. In the absence of any knowledge about the
actual downstream task, we resort to a worst case formulation. An unsupervised
meta-learning algorithm will find a single learning procedure $f$ that has the
lowest regret against an <em>adversarially</em> chosen task distribution $p$:</p>

<p>Our work analyzes how exactly we might obtain such an optimal unsupervised
meta-learner, and provides bounds on the regret that it might incur in the
worst case. Specifically, under some restrictions on the family of tasks that
might be encountered at test-time, the optimal distribution for an unsupervised
meta-learner to propose is <em>uniform</em> over all possible tasks.</p>

<p>The intuition for this is straightforward: if the test time task distribution
can be chosen adversarially, the algorithm must make sure it is uniformly good
over <em>all</em> possible tasks that might be encountered. As a didactic example, if
test-time reward functions were restricted to the class of goal-reaching tasks,
the regret for reaching a goal at test-time is inverse related to the probability
of sampling that goal during training-time. If any one of
the goals $g$ has lower density than the others, an adversary can propose
a task distribution solely consisting of reaching that goal $g$
causing the learning procedure to incur a higher regret. This example suggests that we can find an optimal unsupervised meta-learner using a uniform distribution over goals. Our paper formalizes this idea and extends it to broader classes task distributions.</p>

<p>Now, actually sampling from a uniform distribution over all possible tasks is quite challenging.
Several recent papers have proposed RL exploration methods based on maximizing
mutual information [<a href="https://arxiv.org/abs/1807.10299" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Achiam 2018</a>, <a href="https://arxiv.org/abs/1802.06070" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Eysenbach 2018</a>, <a href="https://arxiv.org/abs/1611.07507" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gregor 2016</a>, <a href="https://arxiv.org/abs/1906.05274" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Lee 2019</a>, <a href="https://arxiv.org/abs/1907.01657" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sharma 2019</a>].
In this work, we show that these methods provide a tractable approximation to the uniform distribution over task distributions. To
understand why this is, we can look at the form of a mutual information
considered by [<a href="https://arxiv.org/abs/1802.06070" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Eysenbach 2018</a>], between states $s$ and latent
variables $z$:</p>

<p>In this objective, the first marginal entropy term is maximized when there is a
uniform distribution over all possible tasks. The second conditional entropy
term ensures consistency, by making sure that for each $z$, the resulting
distribution of $s$ is narrow. This suggests constructing unsupervised
task-distributions in an environment by optimizing mutual information
gives us a provably optimal task distribution, according to our
notion of min-max optimality.</p>

<p>While the analysis makes some limiting assumptions about the forms of tasks
encountered, we show how this analysis can be extended to provide a bound on
the performance in the most general case of reinforcement learning. It also
provides empirical gains on several simulated environments as compared to
methods which train from scratch, as shown in the Figure below.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/umrl/UMRL.png" width=""><br></p>

<h1>Summary &#38; Discussion</h1>

<p>In summary:</p>

<ul><li>
    <p>Learning procedures are recipes for converting datasets into function
approximators. Learning procedures have many knobs, which can be tuned by
optimizing the learning procedures to solve a distribution of tasks.</p>
  </li>
  <li>
    <p>Manually designing these task distributions is challenging, so a recent line
of work suggests that the learning procedure can use unlabeled data to
propose its own tasks for optimizing its knobs.</p>
  </li>
  <li>
    <p>These unsupervised meta-learning algorithms allow for learning in regimes
previously impractical, and further expand that capability of machine
learning methods.</p>
  </li>
  <li>
    <p>This work closely relates to other works on <a href="https://arxiv.org/abs/1907.01657" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">unsupervised</a> <a href="https://arxiv.org/abs/1611.07507" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">skill
discovery</a>, <a href="https://arxiv.org/abs/1705.05363" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">exploration</a> and <a href="https://arxiv.org/abs/1807.03748" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">representation</a> <a href="https://arxiv.org/abs/1511.06434" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">learning</a>, but
explicitly optimizes for transferability of the representations and skills to
downstream tasks.</p>
  </li>
</ul><p>A number of open questions remain about unsupervised meta-learning:</p>

<ul><li>
    <p>Unsupervised learning is closely connected to unsupervised meta-learning: the
former uses unlabeled data to learn features, while the second uses unlabeled
data to tune the learning procedure. Might there be some unifying treatment
of both approaches?</p>
  </li>
  <li>
    <p>Our analysis only proves that task proposal based on mutual information is optimal for memoryless meta-learning algorithms. Meta-learning algorithms with memory, which we expect will perform better, may perform best with different task proposal mechanisms.</p>
  </li>
  <li>
    <p>Scaling unsupervised meta learning to leverage large-scale datasets and
complex tasks holds the promise of acquiring learning procedures for solving
real-world problems more efficiently than our current learning procedures.</p>
  </li>
</ul><p>Check out our paper for more experiments and proofs: <a href="https://arxiv.org/abs/1806.04640" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">https://arxiv.org/abs/1806.04640</a></p>

<h2>Acknowledgments</h2>

<p>Thanks to Jake Tyo, Conor Igoe, Sergey Levine, Chelsea Finn, Misha Khodak, Daniel Seita, and Stefani Karp for their feedback.</p>

<hr><div>
  <ol><li>
      <p>These knobs are often known as hyperparameters, but we will stick with
the colloquial &#8220;knob&#8221; to avoid having to draw a line between parameters and
hyperparameters.&#160;<a href="http://bair.berkeley.edu/blog/#fnref:knob" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">&#8617;</a></p>
    </li>
  </ol></div>]]></description>
										<content:encoded><![CDATA[<p> <strong>By Benjamin Eysenbach and Abhishek Gupta</strong>  </p>
<p><em>This post is cross-listed <a href="https://blog.ml.cmu.edu/2020/05/01/unsupervised-meta-learning-learning-to-learn-without-supervision/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">on the CMU ML blog</a>.</em></p>
<p>The history of machine learning has largely been a story of increasing abstraction. In the dawn of ML, researchers spent considerable effort engineering features. As deep learning gained popularity, researchers then shifted towards tuning the update rules and learning rates for their optimizers. Recent research in meta-learning has climbed one level of abstraction higher: many researchers now spend their days manually constructing task distributions, from which they can automatically learn good optimizers. What might be the next rung on this ladder? In this post we introduce theory and algorithms for <strong>unsupervised meta-learning</strong>, where machine learning algorithms themselves propose their own task distributions. Unsupervised meta-learning further reduces the amount of human supervision required to solve tasks, potentially inserting a new rung on this ladder of abstraction.</p>
<p>  <span id="more-164899"></span>  </p>
<p>We start by discussing how machine learning algorithms use human supervision to find patterns and extract knowledge from observed data. The most common machine learning setting is <em>regression</em>, where a human provides labels $Y$ for a set of examples $X$. The aim is to return a predictor that correctly assigns labels to novel examples. Another common machine learning problem setting is <em>reinforcement learning (RL)</em>, where an agent takes actions in an environment. In RL, humans indicate the desired behavior through a reward function that the agent seeks to maximize. To draw a crude analogy to regression, the environment dynamics are the examples $X$, and the reward function gives the labels $Y$. Algorithms for regression and RL employ many tools, including tabular methods (e.g., value iteration), linear methods (e.g., linear regression) kernel-methods (e.g., RBF-SVMs), and deep neural networks. Broadly, we call these algorithms <em>learning procedures</em>: processes that take as input a dataset (examples with labels, or transitions with rewards) and output a function that performs well (achieves high accuracy or large reward) on the dataset.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/umrl/control_room.jpg" width="500" /> <br /> <i> Machine learning research is similar to the control room for large physics experiments. Researchers have a number of knobs they can tune which affect the performance of the learning procedure. The right setting for the knobs depends on the particular experiment: some settings work well for high-energy experiments; others work well for ultracold atom experiments. <a href="https://www.flickr.com/photos/x-ray_delta_one/3941701730" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Figure Credit</a>. </i> </p>
<p>Similar to lab procedures used in physics and biology, the learning procedures used in machine learning have many knobs<sup id="fnref:knob"><a href="http://bair.berkeley.edu/blog/2020/05/01/umrl/#fn:knob" class="footnote" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">1</a></sup> that can be tuned. For example, the learning procedure for training a neural network might be defined by an optimizer (e.g., Nesterov, Adam) and a learning rate (e.g., 1e-5). Compared with regression, learning procedures specific to RL (e.g., DDPG) often have many more knobs, including the frequency of data collection and how frequently the policy is updated. Finding the right setting for the knobs can have a large effect on how quickly the learning procedure solves a task, and a good configuration of knobs for one learning procedure may be a bad configuration for another.</p>
<h1 id="meta-learning-optimizes-knobs-of-the-learning-procedure">Meta-Learning Optimizes Knobs of the Learning Procedure</h1>
<p>While machine learning practitioners often carefully tune these knobs by hand, if we are going to solve many tasks, it may be useful to automatic this process. The process of setting the knobs of learning procedures via optimization is called <em>meta-learning</em> [<a href="https://link.springer.com/chapter/10.1007/978-1-4615-5529-2_1" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Thrun 1998</a>]. Algorithms that perform this optimization problem automatically are known as <em>meta-learning algorithms</em>. Explicitly tuning the knobs of learning procedures is an active area of research, with various researchers looking at tuning the update rules [<a href="https://arxiv.org/abs/1606.04474" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Andrychowicz 2016</a>, <a href="https://arxiv.org/abs/1611.02779" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Duan 2016</a>, <a href="https://arxiv.org/abs/1611.05763" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wang 2016</a>], weight initialization [<a href="https://arxiv.org/abs/1703.03400" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Finn 2017</a>], network weights [<a href="https://arxiv.org/abs/1609.09106" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Ha 2016</a>], network architectures [<a href="https://arxiv.org/abs/1906.04358" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gaier 2019</a>], and other facets of learning procedures.</p>
<p>To evaluate a setting of knobs, meta-learning algorithms consider not one task but a distribution over many tasks. For example, a distribution over supervised learning tasks may include learning a dog detector, learning a cat detector, and learning a bird detector. In reinforcement learning, a task distribution could be defined as driving a car in a smooth, safe, and efficient manner, where tasks differ by the weights they place on smoothness, safety, and efficiency. Ideally, the task distribution is designed to mirror the distribution over tasks that we are likely to encounter in the real world. Since the tasks in a task distribution are typically related, information from one task may be useful in solving other tasks more efficiently. As you might expect, a knob setting that works best on one distribution of tasks may not be the best for another task distribution; the optimal knob setting depends on the task distribution.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/umrl/Meta-Learning-v2.png" width="" /> <br /> <i> An illustration of meta-learning, where tasks correspond to arranging blocks into different types of towers. The human has a particular block tower in mind and rewards the robot when it builds the correct tower. The robot&#8217;s aim is to build the correct tower as quickly as possible. </i> </p>
<p>In many settings we want to do well on a task distribution to which we have only limited access. For example, in a self-driving car, tasks may correspond to finding the optimal balance of smoothness, safety, and efficiency for each rider, but querying riders to get rewards is expensive. A researcher can attempt to manually construct a task distribution that mimics the true task distribution, but this can be quite challenging and time consuming. Can we avoid having to manually design such task distributions?</p>
<p>To answer this question, we must understand where the benefits of meta-learning come from. When we define task distributions for meta-learning, we do so with some prior knowledge in mind. Without this prior information, tuning the knobs of a learning procedure is often a zero-sum game: setting the knobs to any configuration will accelerate learning on some tasks while slowing learning on other tasks. Does this suggest there is no way to see the benefit of meta-learning without the manual construction of task distributions? Perhaps not! The next section presents an alternative.</p>
<h1 id="optimizing-the-learning-procedure-with-self-proposed-tasks">Optimizing the Learning Procedure with Self-Proposed Tasks</h1>
<p>If designing task distributions is the bottleneck in applying meta-learning algorithms, why not have meta-learning algorithms propose their own tasks? At first glance this seems like a terrible idea, because the No Free Lunch Theorem suggests that this is impossible, <em>without additional knowledge</em>. However, many real-world settings do provide a bit of additional information, albeit disguised as unlabeled data. For example, in regression, we might have access to an unlabeled dataset and know that the downstream tasks will be labeled versions of this same image dataset. In a RL setting, a robot can interact with its environment without receiving any reward, knowing that downstream tasks will be constructed by defining reward functions for this very environment (i.e. the real world). Seen from this perspective, the recipe for <em>unsupervised meta-learning</em> (doing meta-learning without manually constructed tasks) becomes clear: given unlabeled data, construct task distributions from this unlabeled data or environment, and then meta-learn to quickly solve these self-proposed tasks.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/umrl/Unsupervised-Meta-Learning-v2.png" width="" /> <br /> <i> In unsupervised meta-learning, the agent proposes its own tasks, rather than relying on tasks proposed by a human. </i> </p>
<p>How can we use this unlabeled data to construct task distributions which will facilitate learning downstream tasks? In the case of regression, prior work on unsupervised meta-learning [<a href="https://arxiv.org/abs/1810.02334" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Hsu 2018</a>, <a href="https://arxiv.org/abs/1811.11819" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Khodadadeh 2019</a>] clusters an unlabeled dataset of images and then randomly chooses subsets of the clusters to define a distribution of classification tasks. Other work [<a href="https://arxiv.org/abs/1912.04226" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Jabri 2019</a>] look at an RL setting: after exploring an environment without a reward function to collect a set of behaviors that are feasible in this environment, these behaviors are clustered and used to define a distribution of reward functions. In both cases, even though the tasks constructed can be random, the resulting task distribution is not random, because all tasks share the underlying unlabeled data — the image dataset for regression and the environment dynamics for reinforcement learning. <em>The underlying unlabeled data are the inductive bias with which we pay for our free lunch.</em></p>
<p>Let us take a deeper look into the RL case. Without knowing the downstream tasks or reward functions, what is the “best” task distribution for “practicing” to solve tasks quickly? Can we measure how effective a task distribution is for solving unknown, downstream tasks? Is there any sense in which one unsupervised task proposal mechanism is better than another? Understanding the answers to these questions may guide the principled development of meta-learning algorithms with little dependence on human supervision. Our work [<a href="https://arxiv.org/abs/1806.04640" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gupta 2018</a>], takes a first step towards answering these questions. In particular, we examine the <em>worst-case</em> performance of learning procedures, and derive an optimal unsupervised meta-reinforcement learning procedure.</p>
<h1 id="optimal-unsupervised-meta-learning">Optimal Unsupervised Meta-Learning</h1>
<p>To answer the questions posed above, our first step is to define an optimal meta-learner for the case where the distribution of tasks is known. We define an optimal meta-learner as the learning procedure that achieves the largest expected reward, averaged across the distribution of tasks. More precisely, we will compare the expected reward for a learning procedure $f$ to that of best learning procedure $f^*$, defining the <em>regret</em> of $f$ on a task distribution $p$ as follows:</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/umrl/labeled_regret_v2.png" width="600" />  </p>
<p>Extending this definition to the case of unsupervised meta-learning, an optimal unsupervised meta-learner can be defined as a meta-learner that achieves the minimum <em>worst-case</em> regret across all possible task distributions that may be encountered in the environment. In the absence of any knowledge about the actual downstream task, we resort to a worst case formulation. An unsupervised meta-learning algorithm will find a single learning procedure $f$ that has the lowest regret against an <em>adversarially</em> chosen task distribution $p$:</p>
<p>  <script type="math/tex; mode=display">\min_f \max_p \;\; {\rm Regret}(f,p).</script>  </p>
<p>Our work analyzes how exactly we might obtain such an optimal unsupervised meta-learner, and provides bounds on the regret that it might incur in the worst case. Specifically, under some restrictions on the family of tasks that might be encountered at test-time, the optimal distribution for an unsupervised meta-learner to propose is <em>uniform</em> over all possible tasks.</p>
<p>The intuition for this is straightforward: if the test time task distribution can be chosen adversarially, the algorithm must make sure it is uniformly good over <em>all</em> possible tasks that might be encountered. As a didactic example, if test-time reward functions were restricted to the class of goal-reaching tasks, the regret for reaching a goal at test-time is inverse related to the probability of sampling that goal during training-time. If any one of the goals $g$ has lower density than the others, an adversary can propose a task distribution solely consisting of reaching that goal $g$ causing the learning procedure to incur a higher regret. This example suggests that we can find an optimal unsupervised meta-learner using a uniform distribution over goals. Our paper formalizes this idea and extends it to broader classes task distributions.</p>
<p>Now, actually sampling from a uniform distribution over all possible tasks is quite challenging. Several recent papers have proposed RL exploration methods based on maximizing mutual information [<a href="https://arxiv.org/abs/1807.10299" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Achiam 2018</a>, <a href="https://arxiv.org/abs/1802.06070" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Eysenbach 2018</a>, <a href="https://arxiv.org/abs/1611.07507" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gregor 2016</a>, <a href="https://arxiv.org/abs/1906.05274" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Lee 2019</a>, <a href="https://arxiv.org/abs/1907.01657" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sharma 2019</a>]. In this work, we show that these methods provide a tractable approximation to the uniform distribution over task distributions. To understand why this is, we can look at the form of a mutual information considered by [<a href="https://arxiv.org/abs/1802.06070" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Eysenbach 2018</a>], between states $s$ and latent variables $z$:</p>
<p>  <script type="math/tex; mode=display">\mathcal{I}(s,z) = \mathcal{H}(s) - \mathcal{H}(s|z).</script>  </p>
<p>In this objective, the first marginal entropy term is maximized when there is a uniform distribution over all possible tasks. The second conditional entropy term ensures consistency, by making sure that for each $z$, the resulting distribution of $s$ is narrow. This suggests constructing unsupervised task-distributions in an environment by optimizing mutual information gives us a provably optimal task distribution, according to our notion of min-max optimality.</p>
<p>While the analysis makes some limiting assumptions about the forms of tasks encountered, we show how this analysis can be extended to provide a bound on the performance in the most general case of reinforcement learning. It also provides empirical gains on several simulated environments as compared to methods which train from scratch, as shown in the Figure below.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/umrl/UMRL.png" width="" />  </p>
<h1 id="summary--discussion">Summary &amp; Discussion</h1>
<p>In summary:</p>
<ul>
<li>
<p>Learning procedures are recipes for converting datasets into function approximators. Learning procedures have many knobs, which can be tuned by optimizing the learning procedures to solve a distribution of tasks.</p>
</li>
<li>
<p>Manually designing these task distributions is challenging, so a recent line of work suggests that the learning procedure can use unlabeled data to propose its own tasks for optimizing its knobs.</p>
</li>
<li>
<p>These unsupervised meta-learning algorithms allow for learning in regimes previously impractical, and further expand that capability of machine learning methods.</p>
</li>
<li>
<p>This work closely relates to other works on <a href="https://arxiv.org/abs/1907.01657" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">unsupervised</a> <a href="https://arxiv.org/abs/1611.07507" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">skill discovery</a>, <a href="https://arxiv.org/abs/1705.05363" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">exploration</a> and <a href="https://arxiv.org/abs/1807.03748" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">representation</a> <a href="https://arxiv.org/abs/1511.06434" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">learning</a>, but explicitly optimizes for transferability of the representations and skills to downstream tasks.</p>
</li>
</ul>
<p>A number of open questions remain about unsupervised meta-learning:</p>
<ul>
<li>
<p>Unsupervised learning is closely connected to unsupervised meta-learning: the former uses unlabeled data to learn features, while the second uses unlabeled data to tune the learning procedure. Might there be some unifying treatment of both approaches?</p>
</li>
<li>
<p>Our analysis only proves that task proposal based on mutual information is optimal for memoryless meta-learning algorithms. Meta-learning algorithms with memory, which we expect will perform better, may perform best with different task proposal mechanisms.</p>
</li>
<li>
<p>Scaling unsupervised meta learning to leverage large-scale datasets and complex tasks holds the promise of acquiring learning procedures for solving real-world problems more efficiently than our current learning procedures.</p>
</li>
</ul>
<p>Check out our paper for more experiments and proofs: <a href="https://arxiv.org/abs/1806.04640" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">https://arxiv.org/abs/1806.04640</a></p>
<h2 id="acknowledgments">Acknowledgments</h2>
<p>Thanks to Jake Tyo, Conor Igoe, Sergey Levine, Chelsea Finn, Misha Khodak, Daniel Seita, and Stefani Karp for their feedback.This article was initially published on the <a href="https://bair.berkeley.edu/blog/2020/05/01/umrl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
<hr />
<div class="footnotes">
<ol>
<li id="fn:knob">
<p>These knobs are often known as hyperparameters, but we will stick with the colloquial “knob” to avoid having to draw a line between parameters and hyperparameters.&nbsp;<a href="http://bair.berkeley.edu/blog/2020/05/01/umrl/#fnref:knob" class="reversefootnote" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">&#8617;</a></p>
</li>
</ol></div>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Robots learning to move like animals</title>
		<link>https://robohub.org/robots-learning-to-move-like-animals/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Mon, 06 Apr 2020 04:54:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/robots-learning-to-move-like-animals/</guid>

					<description><![CDATA[<!--
TODO TODO TODO
Be careful that these three lines are at the top, and that the title and image change for each blog post!
-->
<p>
<img src="https://bair.berkeley.edu/static/blog/laikago/00_teaser.gif" width=""><br><i>
Quadruped robot learning locomotion skills by imitating a dog.
</i>
</p>

<p>Whether it&#8217;s a dog chasing after a ball, or a monkey swinging through the
trees, animals can effortlessly perform an incredibly rich repertoire of agile
locomotion skills. But designing controllers that enable legged robots to
replicate these agile behaviors can be a very challenging task. The superior
agility seen in animals, as compared to robots, might lead one to wonder: can
we create more agile robotic controllers with less effort by directly imitating
animals?</p>

<p>In this work, we present a framework for learning robotic locomotion skills by
imitating animals. Given a reference motion clip recorded from an animal (e.g.
a dog), our framework uses reinforcement learning to train a control policy
that enables a robot to imitate the motion in the real world. Then, by simply
providing the system with different reference motions, we are able to train a
quadruped robot to perform a diverse set of agile behaviors, ranging from fast
walking gaits to dynamic hops and turns. The policies are trained primarily in
simulation, and then transferred to the real world using a latent space
adaptation technique, which is able to efficiently adapt a policy using only a
few minutes of data from the real robot.</p>

<!--more-->

<div>
  
</div>

<h2>Framework</h2>

<p>Our framework consists of three main components: motion retargeting, motion
imitation, and domain adaptation. 1) First, given a reference motion, the
motion retargeting stage maps the motion from the original animal&#8217;s morphology
to the robot&#8217;s morphology. 2) Next, the motion imitation stage uses the
retargeted reference motion to train a policy for imitating the motion in
simulation. 3) Finally, the domain adaptation stage transfers the policy from
simulation to a real robot via a sample efficient domain adaptation process. We
apply this framework to learn a variety of agile locomotion skills for a
<a href="http://www.unitree.cc/e/action/ShowInfo.php?classid=6&#038;id=1" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Laikago</a>
quadruped robot.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/laikago/01_overview.gif" width=""><br><i>
The framework consists of three stages: motion retargeting, motion imitation,
and domain  adaptation. It receives as input motion data recorded from an
animal, and outputs a control  policy that enables a robot to reproduce the
motion in the real world.
</i>
</p>

<h3>Motion Retargeting</h3>

<p>An animal&#8217;s body is generally quite different from a robot&#8217;s body. So before
the robot can imitate the animal&#8217;s motion, we must first map the motion to the
robot&#8217;s body. The goal of the retargeting process is to construct a reference
motion for the robot that captures the important characteristics of the
animal&#8217;s motion. To do this, we first identify a set of source keypoints on the
animal&#8217;s body, such as the hips and the feet. Then, corresponding target
keypoints are specified on the robot&#8217;s body.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/laikago/02_keypoints_dog.png" height="280"><img src="https://bair.berkeley.edu/static/blog/laikago/02_keypoints_robot.png" height="280"><br><i>
Inverse-kinematics (IK) is used to retarget mocap clips recorded from a real
dog (left) to the robot (right). Corresponding pairs of keypoints (red) are
specified on the dog and robot&#8217;s bodies,  and then IK is used to compute a pose
for the robot that tracks the keypoints.
</i>
</p>

<p>Next, inverse-kinematics is used to construct a reference motion for the robot
that tracks the corresponding keypoints from the animal at every timestep.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/laikago/03_retarget_pace.gif" height="170"><img src="https://bair.berkeley.edu/static/blog/laikago/04_retarget_spin.gif" height="170"><br><i>
Inverse-kinematics is used to retarget mocap clips recorded from a dog to the robot.
</i>
</p>

<h3>Motion Imitation</h3>

<p>After retargeting the reference motion to the robot, the next step is to train
a control policy to imitate the retargeted motion. But reinforcement learning
algorithms can take a long time to learn an effective policy, and directly
training on a real robot can be fairly dangerous (both for the robot and its
human companions). So, we instead opt to perform most of the training in the
comforts of simulation, and then transfer the learned policy to the real world
using more sample efficient adaptation techniques. All simulations are
performed using <a href="https://pybullet.org/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PyBullet</a>.</p>

<p>The policy $\pi(\mathbf{a} &#124; \mathbf{s}, \mathbf{g})$, takes as input a state
$\mathbf{s}$, which represents the configuration of the robot&#8217;s body, and a
goal $\mathbf{g}$, which specifies target poses from the reference motion that
the robot is to imitate.  It then outputs an action $\mathbf{a}$, which
specifies target angles for PD controllers at each of the robot&#8217;s joints. To
train the policy to imitate a reference motion, we use a reward function that
encourages the robot to minimize the difference between the pose of the
reference motion $\hat{\mathbf{q}}_t$ and the pose of the simulated character
$\mathbf{q}_t$ at every timestep $t$,</p>

<p>By simply using different reference motions in the reward function, we can
train a simulated robot to imitate a variety of different skills.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/laikago/05_sim_pace.gif" height="170"><img src="https://bair.berkeley.edu/static/blog/laikago/06_sim_spin.gif" height="170"><br><i>
Reinforcement learning is used to train a simulated robot to imitate the
retargeted reference motions.
</i>
</p>

<h3>Domain Adaptation</h3>

<p>Since simulators generally provide only a coarse approximation of the real
world, policies trained in simulation often perform fairly poorly when deployed
on a real robot. Therefore, to transfer a policy trained in simulation to the
real world, we use a sample efficient domain adaptation techniques that can
adapt the policy to the real world using only a small number of trials on the
real robot. To do this, we first apply
<a href="https://xbpeng.github.io/projects/SimToReal/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">domain randomization</a> during training in simulation, which randomly varies the
dynamics parameters, such as mass and friction. The dynamics parameters are
then also collected into a vector $\mu$ and encoded into a latent presentation
$\mathbf{z}$ by an encoder $E(\mathbf{z} &#124; \mu)$. The latent encoding  is
passed as an additional input to the policy $\pi(\mathbf{a} &#124; \mathbf{s},
\mathbf{g}, \mathbf{z})$.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/laikago/07_policy.png" width="500"><br><i>
The dynamics parameters  of the simulation are varied during training, and also
encoded into a latent representation  that is provided as an additional input
to the policy.
</i>
</p>

<p>When transferring the policy to a real robot, we remove the encoder and
directly search for a $\mathbf{z}$ that maximizes the robot&#8217;s rewards in the
real world.  This is done using
<a href="https://xbpeng.github.io/projects/AWR/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">advantage weighted regression</a>,
a simple off-policy reinforcement learning algorithm. In our experiments, this
technique is often able to adapt a policy to the real world with less than 50
trials, which corresponds to roughly 8 minutes of real-world data.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/laikago/08_adaptation_pace.gif" width=""><br><img src="https://bair.berkeley.edu/static/blog/laikago/09_adaptation_spin.gif" width=""><br><i>
Comparison of policies before and after adaptation on the real robot. Before
adaptation, the robot is prone to falling. But after adaptation, the policies
are able to more consistently execute the desired skills.
</i>
</p>

<h2>Results</h2>

<p>Our framework is able to train a robot to imitate various locomotion skills
from a dog, including different walking gaits, such as pacing and trotting, as
well as a fast spinning motion. By simply playing the forwards walking motions
backwards, we are also able to train the robot to walk backwards.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/laikago/10_real_pace.gif" width="320"><img src="https://bair.berkeley.edu/static/blog/laikago/11_real_trot.gif" width="320"><br><img src="https://bair.berkeley.edu/static/blog/laikago/12_real_spin.gif" width="320"><img src="https://bair.berkeley.edu/static/blog/laikago/13_real_backward_trot.gif" width="320"><br><i>
Laikago imitating various skills from a dog.
</i>
</p>

<p>In addition to imitating motions from real dogs, we can also imitate
artist-animated keyframe motion, including a dynamic hop-turn:</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/laikago/14_real_sidesteps.gif" width="360"><img src="https://bair.berkeley.edu/static/blog/laikago/14_real_turn.gif" width="360"><br>
Hop Turn
<img src="https://bair.berkeley.edu/static/blog/laikago/15_real_hopturn.gif" width=""><br><i>
Skills learned by imitating artist-animated keyframe motions.
</i>
</p>

<p>We also compared the learned policies with the manually-designed controllers
provided by the manufacturer. Our policies are able to learn faster gaits.</p>

<p>
<img src="https://bair.berkeley.edu/static/blog/laikago/16_comp_trot.gif" width=""><br><br><i>
Comparison of learned trotting gait with the built-in gait provided by the
manufacturer.
</i>
</p>

<p>Overall, our system has been able to reproduce a fairly diverse corpus of
behaviors with a quadruped robot. However, due to hardware and algorithmic
limitations, we have not been able to imitate more dynamic motions such as
running and jumping. The learned policies are also not as robust as the best
manually-designed controllers. Exploring techniques for further improving the
agility and robustness of these learned policies could be a valuable step
towards more complex real-world applications. Extending this framework to learn
<a href="https://xbpeng.github.io/projects/SFV/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">skills from videos</a>
would also be an exciting direction, which can substantially increase the
volume of data from which robots can learn from.</p>

<p>To learn more,
<a href="https://xbpeng.github.io/projects/Robotic_Imitation/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">check out the paper and code</a>.</p>

<p>We would like to thank Erwin Coumans, Tingnan Zhang, Tsang-Wei Lee, Jie Tan,
Sergey Levine, Byron David, Thinh Nguyen, Gus Kouretas, Krista Reymann, and
Bonny Ho for all their support and contribution to this work. This project was
done in collaboration with Google Brain.</p>]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/00_teaser.gif" width="900" /><br />   <strong>By Xue Bin (Jason) Peng</strong>  </p>
<p>Whether it’s a dog chasing after a ball, or a monkey swinging through the trees, animals can effortlessly perform an incredibly rich repertoire of agile locomotion skills. But designing controllers that enable legged robots to replicate these agile behaviors can be a very challenging task. The superior agility seen in animals, as compared to robots, might lead one to wonder: can we create more agile robotic controllers with less effort by directly imitating animals?</p>
<p> <span id="more-163545"></span> </p>
<p>In this work, we present a framework for learning robotic locomotion skills by imitating animals. Given a reference motion clip recorded from an animal (e.g. a dog), our framework uses reinforcement learning to train a control policy that enables a robot to imitate the motion in the real world. Then, by simply providing the system with different reference motions, we are able to train a quadruped robot to perform a diverse set of agile behaviors, ranging from fast walking gaits to dynamic hops and turns. The policies are trained primarily in simulation, and then transferred to the real world using a latent space adaptation technique, which is able to efficiently adapt a policy using only a few minutes of data from the real robot.</p>
<div class="keep-aspect"><iframe title="Learning Agile Robotic Locomotion Skills by Imitating Animals" width="500" height="281" src="https://www.youtube-nocookie.com/embed/lKYh6uuCwRY?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></div>
<p></p>
<h2 id="framework">Framework</h2>
<p>Our framework consists of three main components: motion retargeting, motion imitation, and domain adaptation. 1) First, given a reference motion, the motion retargeting stage maps the motion from the original animal’s morphology to the robot’s morphology. 2) Next, the motion imitation stage uses the retargeted reference motion to train a policy for imitating the motion in simulation. 3) Finally, the domain adaptation stage transfers the policy from simulation to a real robot via a sample efficient domain adaptation process. We apply this framework to learn a variety of agile locomotion skills for a <a href="http://www.unitree.cc/e/action/ShowInfo.php?classid=6&amp;id=1" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Laikago</a> quadruped robot.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/01_overview.gif" width="" /><br /> <i> The framework consists of three stages: motion retargeting, motion imitation, and domain  adaptation. It receives as input motion data recorded from an animal, and outputs a control  policy that enables a robot to reproduce the motion in the real world. </i> </p>
<h3 id="motion-retargeting">Motion Retargeting</h3>
<p>An animal’s body is generally quite different from a robot’s body. So before the robot can imitate the animal’s motion, we must first map the motion to the robot’s body. The goal of the retargeting process is to construct a reference motion for the robot that captures the important characteristics of the animal’s motion. To do this, we first identify a set of source keypoints on the animal’s body, such as the hips and the feet. Then, corresponding target keypoints are specified on the robot’s body.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/02_keypoints_dog.png" height="280" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/02_keypoints_robot.png" height="280" /> <br /> <i> Inverse-kinematics (IK) is used to retarget mocap clips recorded from a real dog (left) to the robot (right). Corresponding pairs of keypoints (red) are specified on the dog and robot’s bodies,  and then IK is used to compute a pose for the robot that tracks the keypoints. </i> </p>
<p>Next, inverse-kinematics is used to construct a reference motion for the robot that tracks the corresponding keypoints from the animal at every timestep.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/03_retarget_pace.gif" height="170" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/04_retarget_spin.gif" height="170" /> <br /> <i> Inverse-kinematics is used to retarget mocap clips recorded from a dog to the robot. </i> </p>
<h3 id="motion-imitation">Motion Imitation</h3>
<p>After retargeting the reference motion to the robot, the next step is to train a control policy to imitate the retargeted motion. But reinforcement learning algorithms can take a long time to learn an effective policy, and directly training on a real robot can be fairly dangerous (both for the robot and its human companions). So, we instead opt to perform most of the training in the comforts of simulation, and then transfer the learned policy to the real world using more sample efficient adaptation techniques. All simulations are performed using <a href="https://pybullet.org/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PyBullet</a>.</p>
<p>The policy $\pi(\mathbf{a} | \mathbf{s}, \mathbf{g})$, takes as input a state $\mathbf{s}$, which represents the configuration of the robot’s body, and a goal $\mathbf{g}$, which specifies target poses from the reference motion that the robot is to imitate.  It then outputs an action $\mathbf{a}$, which specifies target angles for PD controllers at each of the robot’s joints. To train the policy to imitate a reference motion, we use a reward function that encourages the robot to minimize the difference between the pose of the reference motion $\hat{\mathbf{q}}_t$ and the pose of the simulated character $\mathbf{q}_t$ at every timestep $t$,</p>
<p>  <script type="math/tex; mode=display">r_t = \exp \Big[ - \| \hat{\mathbf{q}}_t - \mathbf{q}_t \|^2 \Big]</script>  </p>
<p>By simply using different reference motions in the reward function, we can train a simulated robot to imitate a variety of different skills.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/05_sim_pace.gif" height="170" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/06_sim_spin.gif" height="170" /> <br /> <i> Reinforcement learning is used to train a simulated robot to imitate the retargeted reference motions. </i> </p>
<h3 id="domain-adaptation">Domain Adaptation</h3>
<p>Since simulators generally provide only a coarse approximation of the real world, policies trained in simulation often perform fairly poorly when deployed on a real robot. Therefore, to transfer a policy trained in simulation to the real world, we use a sample efficient domain adaptation techniques that can adapt the policy to the real world using only a small number of trials on the real robot. To do this, we first apply <a href="https://xbpeng.github.io/projects/SimToReal/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">domain randomization</a> during training in simulation, which randomly varies the dynamics parameters, such as mass and friction. The dynamics parameters are then also collected into a vector $\mu$ and encoded into a latent presentation $\mathbf{z}$ by an encoder $E(\mathbf{z} | \mu)$. The latent encoding  is passed as an additional input to the policy $\pi(\mathbf{a} | \mathbf{s}, \mathbf{g}, \mathbf{z})$.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/07_policy.png" width="500" /> <br /> <i> The dynamics parameters  of the simulation are varied during training, and also encoded into a latent representation  that is provided as an additional input to the policy. </i> </p>
<p>When transferring the policy to a real robot, we remove the encoder and directly search for a $\mathbf{z}$ that maximizes the robot’s rewards in the real world.  This is done using <a href="https://xbpeng.github.io/projects/AWR/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">advantage weighted regression</a>, a simple off-policy reinforcement learning algorithm. In our experiments, this technique is often able to adapt a policy to the real world with less than 50 trials, which corresponds to roughly 8 minutes of real-world data.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/08_adaptation_pace.gif" width="" /><br /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/09_adaptation_spin.gif" width="" /> <br /> <i> Comparison of policies before and after adaptation on the real robot. Before adaptation, the robot is prone to falling. But after adaptation, the policies are able to more consistently execute the desired skills. </i> </p>
<h2 id="results">Results</h2>
<p>Our framework is able to train a robot to imitate various locomotion skills from a dog, including different walking gaits, such as pacing and trotting, as well as a fast spinning motion. By simply playing the forwards walking motions backwards, we are also able to train the robot to walk backwards.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/10_real_pace.gif" width="320" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/11_real_trot.gif" width="320" /><br /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/12_real_spin.gif" width="320" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/13_real_backward_trot.gif" width="320" /> <br /> <i> Laikago imitating various skills from a dog. </i> </p>
<p>In addition to imitating motions from real dogs, we can also imitate artist-animated keyframe motion, including a dynamic hop-turn:</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/14_real_sidesteps.gif" width="360" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/14_real_turn.gif" width="360" /><br /> Hop Turn <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/15_real_hopturn.gif" width="" /> <br /> <i> Skills learned by imitating artist-animated keyframe motions. </i> </p>
<p>We also compared the learned policies with the manually-designed controllers provided by the manufacturer. Our policies are able to learn faster gaits.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/laikago/16_comp_trot.gif" width="" /></p>
<p> <i> Comparison of learned trotting gait with the built-in gait provided by the manufacturer. </i> </p>
<p>Overall, our system has been able to reproduce a fairly diverse corpus of behaviors with a quadruped robot. However, due to hardware and algorithmic limitations, we have not been able to imitate more dynamic motions such as running and jumping. The learned policies are also not as robust as the best manually-designed controllers. Exploring techniques for further improving the agility and robustness of these learned policies could be a valuable step towards more complex real-world applications. Extending this framework to learn <a href="https://xbpeng.github.io/projects/SFV/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">skills from videos</a> would also be an exciting direction, which can substantially increase the volume of data from which robots can learn from.</p>
<p>To learn more, <a href="https://xbpeng.github.io/projects/Robotic_Imitation/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">check out the paper and code</a>.</p>
<p>We would like to thank Erwin Coumans, Tingnan Zhang, Tsang-Wei Lee, Jie Tan, Sergey Levine, Byron David, Thinh Nguyen, Gus Kouretas, Krista Reymann, and Bonny Ho for all their support and contribution to this work. This project was done in collaboration with Google Brain. This article was initially published on the <a href="https://bair.berkeley.edu/blog/2020/04/03/laikago/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Does on-policy data collection fix errors in off-policy reinforcement learning?</title>
		<link>https://robohub.org/does-on-policy-data-collection-fix-errors-in-off-policy-reinforcement-learning/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 18 Mar 2020 23:29:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/does-on-policy-data-collection-fix-errors-in-off-policy-reinforcement-learning/</guid>

					<description><![CDATA[<!--
TODO TODO TODO
Be careful that these three lines are at the top, and that the title and image change for each blog post!
-->
<p>Reinforcement learning has seen a great deal of success in solving complex decision making problems ranging from <a href="https://arxiv.org/abs/1806.10293" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robotics</a> to <a href="https://deepmind.com/blog/article/AlphaStar-Grandmaster-level-in-StarCraft-II-using-multi-agent-reinforcement-learning" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">games</a> to <a href="http://www.wi-frankfurt.de/publikationenNeu/AReinforcementLearningApproach.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">supply chain management</a> to <a href="https://arxiv.org/pdf/1810.12027.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recommender systems</a>. Despite their success, deep reinforcement learning algorithms can be exceptionally difficult to use, due to unstable training, sensitivity to hyperparameters, and generally unpredictable and poorly understood convergence properties. Multiple explanations, and corresponding solutions, have been proposed for improving the stability of such methods, and we have seen good progress over the last few years on these algorithms. In this blog post, we will dive deep into analyzing a central and underexplored reason behind some of the problems with the class of deep RL algorithms based on dynamic programming, which encompass the popular <a href="https://web.stanford.edu/class/psych209/Readings/MnihEtAlHassibis15NatureControlDeepRL.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DQN</a> and soft actor-critic (<a href="https://arxiv.org/abs/1812.05905" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SAC</a>) algorithms &#8211; the detrimental connection between data distributions and learned models.</p>

&#60;!--
<p style="text-align:center">
<img src="https://paper-attachments.dropbox.com/s_C0A5DF53824B57146C6C7BFA4F136835C682554FAEA4D3AC837E9CAA53C2DDCA_1583955854042_SupvsRL.svg" width="">
<br />
<i>
Figure 1: Distributions can impact the generalization properties of supervised
learning algorithms due to shift between train and test distributions. In RL,
besides generalization, distributions also affects other elements in the
learning process such as the actual updates performed, exploration and has a
significant impact on learning progress even in the absence of explicit
distribution shift.
</i>
</p>
--&#62;

<!--more-->

<p>Before diving deep into a description of this problem, let us quickly recap
some of the main concepts in dynamic programming. Algorithms that apply dynamic
programming in conjunction with function approximation are generally referred
to as approximate dynamic programming (ADP) methods. ADP algorithms include
some of the most popular, state-of-the-art RL methods such as variants of deep
Q-networks (DQN) and soft actor-critic (SAC) algorithms. ADP methods based on
Q-learning train action-value functions, , via a Bellman backup. In
practice, this corresponds to training a parametric function, , by minimizing the mean squared difference to a backup estimate of the
Q-function, defined as:</p>

<p>where  denotes a previous instance of the original Q-function,
, and is commonly referred to as a target network. This update is
summarized in the equation below.</p>

<p>An analogous update is also used for
<a href="https://papers.nips.cc/paper/1786-actor-critic-algorithms.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">actor-critic</a>
methods that also maintain an explicitly parametrized policy,
, alongside a Q-function. Such an update typically replaces
 with an expectation under the policy, . We shall use the  version for consistency throughout,
however, the actor-critic version follows analogously. These ADP methods aim at
learning the optimal value function, , by applying the Bellman backup
iteratively untill convergence.</p>

<!--
***
-->

<p>A central factor that affects the performance of ADP algorithms is the choice
of the training data-distribution, , as shown in the equation
above. The choice of  is an integral component of the backup,
and it affects solutions obtained via ADP methods, especially since function
approximation is involved. Unlike tabular settings, function approximation
causes the learned Q function to depend on the choice of data distribution
, thereby affecting the dynamics of the learning process. We
show that on-policy exploration induces distributions  such that
training Q-functions under  may fail to correct systematic
errors in the Q-function, even if Bellman error is minimized as much as
possible &#8211; a phenomenon that we refer to as an absence of <strong><em>corrective
feedback</em></strong>.</p>

<h1>Corrective Feedback and Why it is Absent in ADP</h1>

<p>What is corrective feedback formally? How do we determine if it is present or
absent in ADP methods? In order to build intuition, we first present a simple
contextual bandit (one step RL) example, where the Q-function is trained to
match  via supervised updates, without bootstrapping. This enjoys
corrective feedback, and we then contrast it with ADP methods, which do not. In
this example, the goal is to learn the optimal value function ,
which, is equal to the reward . At iteration , the algorithm
minimizes the estimation error of the Q-function:</p>

<p>Using an -greedy or Boltzmann policy for exploration, denoted by $\pi_k$, gives rise
to a <em>hard negative mining</em> phenomenon &#8211; the policy chooses precisely those
actions that correspond to possibly over-estimated Q-values for each state
 and observes the corresponding,  or , as a
result.  Then, minimizing , on samples collected this way
corrects errors in the Q-function, as  is pushed closer to match
 for actions  with incorrectly high Q-values, correcting
precisely the Q-values which may cause sub-optimal performance. This
constructive interaction between online data collection and error correction &#8211;
where the induced online data distribution <em>corrects</em> errors in the value
function &#8211; is what we refer to as <strong>corrective feedback</strong>.</p>

<p>In contrast, we will demonstrate that ADP methods that rely on previous
Q-functions to generate targets for training the current Q-function, may not
benefit from corrective feedback. This difference between bandits and ADP
happens because the target values are computed by applying a  Bellman backup on
the previous Q-function,  (target value), rather than the optimal
, so, errors in , at the next states can result in incorrect
Q-value targets at the current state. No matter how often the current
transition is observed, or how accurately Bellman errors are minimized, the
error in the Q-value with respect to the optimal Q-function, , at
this state is not reduced. Furthermore, in order to obtain correct target
values, we need to ensure that values at state-action pairs occurring at the
tail ends of the data distribution , which are primary causes of
errors in Q-values at other states, are correct. However, as we will show via a
simple didactic example, that this correction process may be extremely slow and
may not occur, mainly because of undesirable generalization effects of the
function approximator.</p>

<p>Let&#8217;s consider a didactic example of a tree-structured deterministic MDP with 7
states and 2 actions,  and , at each state.</p>

<p>
<img src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1584253706874_on_policy_figure_aliasing.png" width=""><br><i>
Figure 1: Run of an ADP algorithm with on-policy data collection. Boxed nodes
and circled nodes denote groups of states aliased by function approximation --
values of these nodes are affected due to parameter sharing and function
approximation.
</i>
</p>

<p>A run of an ADP algorithm that chooses the current on-policy state-action
marginal as  on this tree MDP is shown in Figure 1.  Thus,
the Bellman error at a state is minimized in proportion to the frequency of
occurrence of that state in the policy state-action marginal. Since the leaf
node states are the least frequent in this on-policy marginal distribution (due
to the discounting), the Bellman backup is unable to correct errors in Q-values
at such leaf nodes, due to their low frequency and aliasing with other states
arising due to function approximation. Using incorrect Q-values at the leaf
nodes to generate targets for other nodes in the tree, just gives rise to
incorrect values, even if Bellman error is fully minimized at those states.
Thus, most of the Bellman updates do not actually bring Q-values at the states
of the MDP closer to , since the primary cause of incorrect target
values isn&#8217;t corrected.</p>

<p>This observation is surprising, since it demonstrates how the choice of an
online distribution coupled with function approximation might actually learn
incorrect Q-values. On the other hand, a scheme that chooses to update states
level by level progressively (Figure 2), ensuring that target values used at
any iteration of learning are correct, very easily learns correct Q-values in
this example.</p>

<!--
![Figure 2: Run of an ADP algorithm with an oracle distribution, that updates states level-by level, progressing through the tree from the leaves to the root. Even in the presence of function approximation, selecting the right set of nodes for updates gives rise to correct Q-values.](https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1584341988034_discor_func_approx_final.png)
-->

<p>
<img src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1584341988034_discor_func_approx_final.png" width=""><br><i>
Figure 2:  Run of an ADP algorithm with an oracle distribution, that updates
states level-by level, progressing through the tree from the leaves to the
root. Even in the presence of function approximation, selecting the right set
of nodes for updates gives rise to correct Q-values.
</i>
</p>

<h1>Consequences of Absent Corrective Feedback</h1>

<p>Now, one might ask if an absence of corrective feedback occurs in practice,
beyond a simple didactic example and whether it hurts in practical problems.
Since visualizing the dynamics of the learning process is hard in practical
problems as we did for the didactic example, we instead devise a metric that
quantifies our intuition for corrective feedback. This metric, what we call
<em>value error,</em> is given by:</p>

<p>Increasing values of  imply that the algorithm is pushing
Q-values farther away from , which means that corrective feedback is
absent, if this happens over a number of iterations. On the other hand,
decreasing values of  implies that the algorithm is
continuously improving its estimate of , by moving it towards  with
each iteration, indicating the presence of corrective feedback.</p>

<p>Observe in Figure 3, that ADP methods can suffer from prolonged periods where
this global measure of error in the Q-function, , is
increasing or fluctuating, and the corresponding returns degrade or stagnate,
implying an absence of corrective feedback.</p>

<!--
![Figure 3: Consequences of absent corrective feedback, including (a) sub-optimal convergence, (b) instability in learning and (c) inability to learn with sparse rewards.](https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1583775913964_Screen+Shot+2020-03-09+at+10.44.58+AM.png)
-->

<p>
<img src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1583775913964_Screen+Shot+2020-03-09+at+10.44.58+AM.png" width=""><br><i>
Figure 3: Consequences of absent corrective feedback, including (a) sub-optimal
convergence, (b) instability in learning and (c) inability to learn with sparse
rewards.
</i>
</p>

<p>In particular, we describe three different consequences of an absence of
corrective feedback:</p>

<ol><li>
    <p><strong>Convergence to suboptimal Q-functions.</strong> We find that on-policy sampling
can cause ADP to converge to a suboptimal solution, even in the absence of
sampling error. Figure 3(a) shows that the value error 
rapidly decreases initially, and eventually converges to a value significantly
greater than 0, from which the learning process never recovers.</p>
  </li>
  <li>
    <p><strong>Instability in the learning process.</strong> We observe that ADP with replay
buffers can be unstable. For instance, the algorithm is prone to degradation
even if the latest policy obtains returns that are very close to the optimal
return in Figure 3(b).</p>
  </li>
  <li>
    <p><strong>Inability to learn with low signal-to-noise ratio.</strong> Absence of corrective
feedback can also prevent ADP algorithms from learning quickly in scenarios
with low signal-to-noise ratio, such as tasks with sparse/noisy rewards as
shown in Figure 3(c). Note that this is not an exploration issue, since all
transitions in the MDP are provided to the algorithm in this experiment.</p>
  </li>
</ol><h1>Inducing Maximal Corrective Feedback via Distribution Correction</h1>

<p>Now that we have defined corrective feedback and gone over some detrimental
consequences an absence of it can have on the learning process of an ADP
algorithm, what might be some ways to fix this problem? To recap, an absence of
corrective feedback occurs when ADP algorithms naively use the on-policy or
replay buffer distributions for training Q-functions. One way to prevent this
problem is by computing an &#8220;optimal&#8221; data distribution that provides maximal
corrective feedback, and train Q-functions using this distribution? This way we
can ensure that the ADP algorithm always enjoys corrective feedback, and hence
makes steady learning progress. The strategy we used in our work is to compute
this optimal distribution and then perform a <strong>weighted Bellman update</strong> that
re-weights the data distribution in the replay buffer to this optimal
distribution (in practice, a tractable approximation is required, as we will
see) via importance sampling based techniques.</p>

<p>We will not go into the full details of our derivation in this article,
however, we mention the optimization problem used to obtain a form for this
optimal distribution and encourage readers interested in the theory to checkout
Section 4 in our paper.  In this optimization problem, our goal is to minimize
a measure of corrective feedback, given by <em>value</em> error ,
with respect to the distribution  used for Bellman error minimization,
at every iteration . This gives rise to the following problem:</p>

<p>We show in our paper that the solution of this optimization problem, that we
refer to as the optimal distribution, , is given by:</p>

<p>By simplifying this expression, we obtain a practically viable expression for
weights, , at any iteration  that can be used to re-weight the data
distribution:</p>

<p>where  is the accumulated Bellman error over iterations, and it
satisfies a convenient recursion making it amenable to practical
implementations,</p>

<p>and  is the Boltzmann or greedy policy
corresponding to the current Q-function.</p>

<p>What does this expression for  intuitively correspond to? Observe that
the term appearing in the exponent in the expression for  corresponds to
the accumulated Bellman error in the target values. Our choice of ,
thus, basically down-weights transitions with highly incorrect target values.
This technique falls into a broader class of <strong>abstention</strong> based techniques
that are common in supervised learning settings with noisy labels, where
down-weighting datapoints (transitions here) with errorful labels (target
values here) can boost generalization and correctness properties of the learned
model.</p>

<!--
![Figure 4: Schematic of the DisCor algorithm. Transitions with errorful target values are downweighted.](https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1584342258383_discor_schematic.png)
-->

<p>
<img src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1584342258383_discor_schematic.png" width=""><br><i>
Figure 4: Schematic of the DisCor algorithm. Transitions with errorful target
values are downweighted.
</i>
</p>

<p>Why does our choice of , i.e. the sum of accumulated Bellman errors
suffice? This is because this value  accounts for how error is
propagated in ADP methods. Bellman errors,  are
propagated under the current policy , and then discounted when
computing target values for updates in ADP.  captures exactly this,
and therefore, using this estimate in our weights suffices.</p>

<p>Our practical algorithm, that we refer to as <strong>D</strong>is<strong>C</strong>or (Distribution
Correction), is identical to conventional ADP methods like Q-learning, with the
exception that it performs a weighted Bellman backup &#8211; it assigns a weight
 to a transition,  and performs a Bellman backup
weighted by these weights, as shown below.</p>

<p>We depict the general principle in the schematic diagram shown in Figure 4.</p>

<h1>How does DisCor perform in practice?</h1>

<p>We finally present some results that demonstrate the efficacy of our method,
DisCor, in practical scenarios. Since DisCor only modifies the chosen
distribution for the Bellman update, it can be applied on top of any standard
ADP algorithm including soft actor-critic (SAC) or deep Q-network (DQN). Our
paper presents results for a number of tasks spanning a wide variety of
settings including robotic manipulation tasks, multi-task reinforcement
learning tasks, learning with stochastic and noisy rewards, and Atari games.
In this blog post, we present two of these results from robotic manipulation
and multi-task RL.</p>

<ol><li>
    <p><strong>Robotic manipulation tasks.</strong> On six challenging benchmark tasks from the
<a href="https://meta-world.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">MetaWorld</a> suite, we observe that DisCor
when combined with SAC greatly outperforms prior state-of-the-art RL
algorithms such as soft actor-critic (SAC) and prioritized experience replay
(<a href="https://arxiv.org/abs/1511.05952" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PER</a>) which is a prior method that
prioritizes states with high Bellman error during training. Note that DisCor
usually starts learning earlier than other methods compared to. DisCor
outperforms vanilla SAC by a factor of about <strong>50%</strong> on average, in terms of
success rate on these tasks.</p>

    <p>
<img src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1583776003015_Screen+Shot+2020-03-09+at+10.45.54+AM.png" width=""><br></p>
  </li>
  <li>
    <p><strong>Multi-task reinforcement learning.</strong> We also present certain results on
the Multi-task 10 (MT10) and Multi-task 50 (MT50) benchmarks from the
Meta-world suite. The goal here is to learn a single policy that can solve a
number of (10 or 50, respectively) different manipulation tasks that share
common structure.  We note that DisCor outperforms, state-of-the-art SAC
algorithm on both of these benchmarks by a wide margin (for e.g. <strong>50%</strong> on
MT10, success rate). Unlike the learning process of SAC that tends to
plateau over the course of learning, we observe that DisCor always exhibits
a non-zero gradient for the learning process, until it converges.</p>

    <p>
<img src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1584342410332_Screen+Shot+2020-03-16+at+12.06.22+AM.png" width=""><br></p>
  </li>
</ol><!--
*
--><p>In our paper, we also perform evaluations on other domains such as Atari games
and OpenAI gym benchmarks, and we encourage the readers to check those out. We
also perform an analysis of the method on tabular domains, understanding
different aspects of the method.</p>

<h1>Perspectives, Future Work and Open Problems</h1>

<p>Some of <a href="https://arxiv.org/abs/1906.00949" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our</a> and
<a href="https://arxiv.org/abs/1712.06924" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">other</a> prior work has highlighted the impact
of the data distribution on the performance of ADP algorithms, We observed in
another <a href="https://arxiv.org/abs/1902.10250" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">prior work</a> that in contrast to the
intuitive belief about the efficacy of online Q-learning with on-policy data
collection, Q-learning with a uniform distribution over states and actions
seemed to perform best. Obtaining a uniform distribution over state-action
tuples during training is not possible in RL, unless all states and actions are
observed at least once, which may not be the case in a number of scenarios. We
might also ask the question about whether the uniform distribution is the best
choice that can be used in an RL setting? The form of the optimal distribution
derived in Section 4 of our paper, is a potentially better choice since it is
customized to the MDP under consideration.</p>

<p>Furthermore, in the domain of purely offline reinforcement learning, studied in
<a href="https://arxiv.org/abs/1906.00949" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our</a> prior work and some other works, such
as <a href="https://arxiv.org/abs/1712.06924" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this</a> and
<a href="https://arxiv.org/abs/1907.04543" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this</a>, we observe that the data distribution
is again a central feature, where backing up out-of-distribution actions and
the inability to try these actions out in the environment to obtain answers to
counterfactual queries, can cause error accumulation and backups to diverge.
However, in this work, we demonstrate a somewhat counterintuitive finding: even
with on-policy data collection, where the algorithm, in principle, can evaluate
all forms of counterfactual queries, the algorithm may not obtain a steady
learning progress, due to an undesirable interaction between the data
distribution and generalization effects of the function approximator.</p>

<p>
<img src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1583852356782_batchgraph.png" width=""><br></p>

<h2>What might be a few promising directions to pursue in future work?</h2>

<p><strong>Formal analysis of learning dynamics:</strong> While our study is an initial foray
into the role that data distributions play in the learning dynamics of ADP
algorithms, this motivates a significantly deeper direction of future study. We
need to answer questions related to how deep neural network based function
approximators actually behave, which are behind these ADP methods, in order to
get them to enjoy corrective feedback.</p>

<p><strong>Re-weighting to supplement exploration in RL problems:</strong> Our work depicts the
promise of re-weighting techniques as a practically simple replacement for
altering entire exploration strategies. We believe that re-weighting techniques
are very promising as a general tool in our toolkit to develop RL algorithms.
In an online RL setting, re-weighting can help remove the some of the burden
off exploration algorithms, and can thus, potentially help us employ complex
exploration strategies in RL algorithms.</p>

<p>More generally, we would like to make a case of analyzing effects of data
distribution more deeply in the context of deep RL algorithms. It is well known
that narrow distributions can lead to brittle solutions in supervised learning
that also do not generalize. What is the corresponding analogue in
reinforcement learning? Distributional robustness style techniques have been
used in supervised learning to guarantee a uniformly convergent learning
process, but it still remains unclear how to apply these in an RL with function
approximation setting. Part of the reason is that the theory of RL often
derives from tabular settings, where distributions do not hamper the learning
process to the extent they do with function approximation. However, as we
showed in this work, choosing the right distribution may lead to significant
gains in deep RL methods, and therefore, we believe, that this issue should be
studied in more detail.</p>

<hr><p>This blog post is based on our recent paper:</p>

<ul><li><strong>DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction</strong> <br>
Aviral Kumar, Abhishek Gupta, Sergey Levine <br><a href="https://arxiv.org/abs/2003.07305" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">arXiv</a></li>
</ul><p>We thank Sergey Levine and Marvin Zhang for their valuable feedback on this blog post.</p>]]></description>
										<content:encoded><![CDATA[<img decoding="async" src="https://aihub.org/wp-content/uploads/2020/03/DisCor.png" alt="" width="915" height="358" class="alignnone size-full wp-image-1764" />
<p>Reinforcement learning has seen a great deal of success in solving complex decision making problems ranging from <a href="https://arxiv.org/abs/1806.10293" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robotics</a> to <a href="https://deepmind.com/blog/article/AlphaStar-Grandmaster-level-in-StarCraft-II-using-multi-agent reinforcement-learning" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">games</a> to <a href="http://www.wi frankfurt.de/publikationenNeu/AReinforcementLearningApproach.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">supply chain management</a> to <a href="https://arxiv.org/pdf/1810.12027.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recommender systems</a>. Despite their success, deep reinforcement learning algorithms can be exceptionally difficult to use, due to unstable training, sensitivity to hyperparameters, and generally unpredictable and poorly understood convergence properties. Multiple explanations, and corresponding solutions, have been proposed for improving the stability of such methods, and we have seen good progress over the last few years on these algorithms. In this blog post, we will dive deep into analyzing a central and underexplored reason behind some of the problems with the class of deep RL algorithms based on dynamic programming, which encompass the popular <a href="https://web.stanford.edu/class/psych209/Readings/MnihEtAlHassibis15NatureControlDeepRL.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DQN<a> and soft actor-critic (<a href="https://arxiv.org/abs/1812.05905">SAC</a>) algorithms – the detrimental connection between data distributions and learned models.</p>
<p><!--


<p style="text-align:center;">
<img decoding="async" src="https://paper-attachments.dropbox.com/s_C0A5DF53824B57146C6C7BFA4F136835C682554FAEA4D3AC837E9CAA53C2DDCA_1583955854042_SupvsRL.svg" width="">
<br />
<i>
Figure 1: Distributions can impact the generalization properties of supervised learning algorithms due to shift between train and test distributions. In RL, besides generalization, distributions also affects other elements in the learning process such as the actual updates performed, exploration and has a significant impact on learning progress even in the absence of explicit distribution shift.
</i>
</p>


--></p>
<p><span id="more-161664"></span></p>
<p>Before diving deep into a description of this problem, let us quickly recap some of the main concepts in dynamic programming. Algorithms that apply dynamic programming in conjunction with function approximation are generally referred to as approximate dynamic programming (ADP) methods. ADP algorithms include some of the most popular, state-of-the-art RL methods such as variants of deep Q-networks (DQN) and soft actor-critic (SAC) algorithms. ADP methods based on Q-learning train action-value functions, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-65390a3fba1f4491907928189cbe07b0_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#40;&#115;&#44;&#32;&#97;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="52" style="vertical-align: -5px;"/>, via a Bellman backup. In practice, this corresponds to training a parametric function, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-4463f8ffaa3ad7c66ca64ea75b02a122_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#95;&#92;&#116;&#104;&#101;&#116;&#97;&#40;&#115;&#44;&#32;&#97;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="60" style="vertical-align: -5px;"/>, by minimizing the mean squared difference to a backup estimate of the Q-function, defined as:</p>
<p><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-921bb7bc2d3ee14150b6bb7b6f2e77be_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#66;&#125;&#94;&#42;&#81;&#40;&#115;&#44;&#32;&#97;&#41;&#32;&#61;&#32;&#114;&#40;&#115;&#44;&#32;&#97;&#41;&#32;&#43;&#32;&#92;&#103;&#97;&#109;&#109;&#97;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#69;&#125;&#95;&#123;&#115;&#39;&#124;&#115;&#44;&#32;&#97;&#125;&#32;&#91;&#92;&#109;&#97;&#120;&#95;&#123;&#97;&#39;&#125;&#32;&#92;&#98;&#97;&#114;&#123;&#81;&#125;&#40;&#115;&#39;&#44;&#32;&#97;&#39;&#41;&#93;&#44;" title="Rendered by QuickLaTeX.com" height="23" width="346" style="vertical-align: -8px;"/></p>
<p>where <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-da235bcb676c5a20152a158cd137f644_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#98;&#97;&#114;&#123;&#81;&#125;" title="Rendered by QuickLaTeX.com" height="19" width="14" style="vertical-align: -4px;"/> denotes a previous instance of the original Q-function, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-79e7881605f06d69cf848e66d5f7350b_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#95;&#92;&#116;&#104;&#101;&#116;&#97;" title="Rendered by QuickLaTeX.com" height="16" width="21" style="vertical-align: -4px;"/>, and is commonly referred to as a target network. This update is summarized in the equation below.</p>
<p><script type="math/tex; mode=display">\theta \leftarrow \arg \min_\theta \mathbb{E}_{s, a \sim \mathcal{D}} \left[(Q_\theta(s, a)- (r(s, a) + \gamma \mathbb{E}_{s'|s, a} [\max_{a'}\bar{Q}(s', a')]))^2 \right]</script></p>
<p>An analogous update is also used for <a href="https://papers.nips.cc/paper/1786-actor-critic algorithms.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">actor-critic</a> methods that also maintain an explicitly parametrized policy, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-d6a1215446046847c719d604630ed64c_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#112;&#105;&#95;&#92;&#112;&#104;&#105;&#40;&#97;&#124;&#115;&#41;" title="Rendered by QuickLaTeX.com" height="20" width="54" style="vertical-align: -6px;"/>, alongside a Q-function. Such an update typically replaces <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-5e19c429dff382e3fd36bd4ee7a19224_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#120;&#95;&#123;&#97;&#39;&#125;" title="Rendered by QuickLaTeX.com" height="11" width="44" style="vertical-align: -3px;"/> with an expectation under the policy, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-c2a5d0588a905dd3433a2c0b3e424ba8_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#69;&#125;&#95;&#123;&#97;&#39;&#32;&#92;&#115;&#105;&#109;&#32;&#92;&#112;&#105;&#95;&#92;&#112;&#104;&#105;&#125;" title="Rendered by QuickLaTeX.com" height="19" width="49" style="vertical-align: -7px;"/>. We shall use the <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-5e19c429dff382e3fd36bd4ee7a19224_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#120;&#95;&#123;&#97;&#39;&#125;" title="Rendered by QuickLaTeX.com" height="11" width="44" style="vertical-align: -3px;"/> version for consistency throughout, however, the actor-critic version follows analogously. These ADP methods aim at learning the optimal value function, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-285dd7f4fb9e9c4e2f9bc31308771a4e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#94;&#42;" title="Rendered by QuickLaTeX.com" height="16" width="20" style="vertical-align: -4px;"/>, by applying the Bellman backup iteratively untill convergence.</p>
<p><!--
***
--></p>
<p>A central factor that affects the performance of ADP algorithms is the choice of the training data-distribution, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-604919a910da61f809c93c64c0740a83_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#68;&#125;" title="Rendered by QuickLaTeX.com" height="12" width="14" style="vertical-align: 0px;"/>, as shown in the equation above. The choice of <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-604919a910da61f809c93c64c0740a83_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#68;&#125;" title="Rendered by QuickLaTeX.com" height="12" width="14" style="vertical-align: 0px;"/> is an integral component of the backup, and it affects solutions obtained via ADP methods, especially since function approximation is involved. Unlike tabular settings, function approximation causes the learned Q function to depend on the choice of data distribution <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-604919a910da61f809c93c64c0740a83_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#68;&#125;" title="Rendered by QuickLaTeX.com" height="12" width="14" style="vertical-align: 0px;"/>, thereby affecting the dynamics of the learning process. We show that on-policy exploration induces distributions <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-604919a910da61f809c93c64c0740a83_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#68;&#125;" title="Rendered by QuickLaTeX.com" height="12" width="14" style="vertical-align: 0px;"/> such that training Q-functions under <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-604919a910da61f809c93c64c0740a83_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#68;&#125;" title="Rendered by QuickLaTeX.com" height="12" width="14" style="vertical-align: 0px;"/> may fail to correct systematic errors in the Q-function, even if Bellman error is minimized as much as possible – a phenomenon that we refer to as an absence of <strong><em>corrective feedback</em></strong>.</p>
<h1 id="corrective-feedback-and-why-it-is-absent-in-adp">Corrective Feedback and Why it is Absent in ADP</h1>
<p>What is corrective feedback formally? How do we determine if it is present or absent in ADP methods? In order to build intuition, we first present a simple contextual bandit (one step RL) example, where the Q-function is trained to match <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-285dd7f4fb9e9c4e2f9bc31308771a4e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#94;&#42;" title="Rendered by QuickLaTeX.com" height="16" width="20" style="vertical-align: -4px;"/> via supervised updates, without bootstrapping. This enjoys corrective feedback, and we then contrast it with ADP methods, which do not. In this example, the goal is to learn the optimal value function <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-e385fc99f892aa40814ec959316c5b89_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#94;&#42;&#40;&#115;&#44;&#32;&#97;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="60" style="vertical-align: -5px;"/>, which, is equal to the reward <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-fdd649aa12813f79e6432efeac936174_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#114;&#40;&#115;&#44;&#32;&#97;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="47" style="vertical-align: -5px;"/>. At iteration <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3422b6bb5c160593658b7c39425d9880_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#107;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: 0px;"/>, the algorithm minimizes the estimation error of the Q-function:</p>
<p><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-e00b957bd30a25554c14f209fd5836b4_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#40;&#81;&#41;&#32;&#61;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#69;&#125;&#95;&#123;&#115;&#32;&#92;&#115;&#105;&#109;&#32;&#92;&#98;&#101;&#116;&#97;&#40;&#115;&#41;&#44;&#32;&#97;&#32;&#92;&#115;&#105;&#109;&#32;&#92;&#112;&#105;&#95;&#107;&#40;&#97;&#124;&#115;&#41;&#125;&#91;&#124;&#81;&#95;&#107;&#40;&#115;&#44;&#32;&#97;&#41;&#32;&#45;&#32;&#81;&#94;&#42;&#40;&#115;&#44;&#32;&#97;&#41;&#124;&#93;&#46;" title="Rendered by QuickLaTeX.com" height="22" width="352" style="vertical-align: -8px;"/></p>
<p>Using an <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ecb3ec0fc6866e26ec884145e44be817_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#118;&#97;&#114;&#101;&#112;&#115;&#105;&#108;&#111;&#110;" title="Rendered by QuickLaTeX.com" height="8" width="8" style="vertical-align: 0px;"/>-greedy or Boltzmann policy for exploration, denoted by <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-88d17dd5f52992a623a09d7bd3fa4be5_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#112;&#105;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="11" width="17" style="vertical-align: -3px;"/>, gives rise to a <em>hard negative mining</em> phenomenon – the policy chooses precisely those actions that correspond to possibly over-estimated Q-values for each state <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ae1901659f469e6be883797bfd30f4f8_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#115;" title="Rendered by QuickLaTeX.com" height="8" width="8" style="vertical-align: 0px;"/> and observes the corresponding, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-fdd649aa12813f79e6432efeac936174_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#114;&#40;&#115;&#44;&#32;&#97;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="47" style="vertical-align: -5px;"/> or <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-e385fc99f892aa40814ec959316c5b89_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#94;&#42;&#40;&#115;&#44;&#32;&#97;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="60" style="vertical-align: -5px;"/>, as a result.  Then, minimizing <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-b257ec690f31847d01cc43d85785a7e6_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#76;&#125;&#40;&#81;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="39" style="vertical-align: -5px;"/>, on samples collected this way corrects errors in the Q-function, as <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-9ee62f9397ef9d893504c9ba837f8244_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#95;&#107;&#40;&#115;&#44;&#32;&#97;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="60" style="vertical-align: -5px;"/> is pushed closer to match <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-e385fc99f892aa40814ec959316c5b89_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#94;&#42;&#40;&#115;&#44;&#32;&#97;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="60" style="vertical-align: -5px;"/> for actions <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-5c53d6ebabdbcfa4e107550ea60b1b19_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#97;" title="Rendered by QuickLaTeX.com" height="8" width="9" style="vertical-align: 0px;"/> with incorrectly high Q-values, correcting precisely the Q-values which may cause sub-optimal performance. This constructive interaction between online data collection and error correction – where the induced online data distribution <em>corrects</em> errors in the value function – is what we refer to as <strong>corrective feedback</strong>.</p>
<p>In contrast, we will demonstrate that ADP methods that rely on previous Q-functions to generate targets for training the current Q-function, may not benefit from corrective feedback. This difference between bandits and ADP happens because the target values are computed by applying a  Bellman backup on the previous Q-function, <script type="math/tex">\bar{Q}</script> (target value), rather than the optimal <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-285dd7f4fb9e9c4e2f9bc31308771a4e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#94;&#42;" title="Rendered by QuickLaTeX.com" height="16" width="20" style="vertical-align: -4px;"/>, so, errors in <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-da235bcb676c5a20152a158cd137f644_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#98;&#97;&#114;&#123;&#81;&#125;" title="Rendered by QuickLaTeX.com" height="19" width="14" style="vertical-align: -4px;"/>, at the next states can result in incorrect Q-value targets at the current state. No matter how often the current transition is observed, or how accurately Bellman errors are minimized, the error in the Q-value with respect to the optimal Q-function, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-54951895d295096dd2f2bd252eb0aa84_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#124;&#81;&#32;&#45;&#32;&#81;&#94;&#42;&#124;" title="Rendered by QuickLaTeX.com" height="19" width="63" style="vertical-align: -5px;"/>, at this state is not reduced. Furthermore, in order to obtain correct target values, we need to ensure that values at state-action pairs occurring at the tail ends of the data distribution <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-604919a910da61f809c93c64c0740a83_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#68;&#125;" title="Rendered by QuickLaTeX.com" height="12" width="14" style="vertical-align: 0px;"/>, which are primary causes of errors in Q-values at other states, are correct. However, as we will show via a simple didactic example, that this correction process may be extremely slow and may not occur, mainly because of undesirable generalization effects of the function approximator.</p>
<p>Let’s consider a didactic example of a tree-structured deterministic MDP with 7 states and 2 actions, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-460f5abc45b558edf34fe288dc0a9979_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#97;&#95;&#49;" title="Rendered by QuickLaTeX.com" height="11" width="15" style="vertical-align: -3px;"/> and <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-0429f295f30eec4db774f087e7f85c21_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#97;&#95;&#50;" title="Rendered by QuickLaTeX.com" height="11" width="16" style="vertical-align: -3px;"/>, at each state.</p>
<p style="text-align:center;">
<img decoding="async" src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1584253706874_on_policy_figure_aliasing.png" width="" /><br />
<br />
<i><br />
Figure 1: Run of an ADP algorithm with on-policy data collection. Boxed nodes and circled nodes denote groups of states aliased by function approximation &#8212; values of these nodes are affected due to parameter sharing and function approximation.<br />
</i>
</p>
<p>A run of an ADP algorithm that chooses the current on-policy state-action marginal as <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-604919a910da61f809c93c64c0740a83_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#68;&#125;" title="Rendered by QuickLaTeX.com" height="12" width="14" style="vertical-align: 0px;"/> on this tree MDP is shown in Figure 1.  Thus, the Bellman error at a state is minimized in proportion to the frequency of occurrence of that state in the policy state-action marginal. Since the leaf node states are the least frequent in this on-policy marginal distribution (due to the discounting), the Bellman backup is unable to correct errors in Q-values at such leaf nodes, due to their low frequency and aliasing with other states arising due to function approximation. Using incorrect Q-values at the leaf nodes to generate targets for other nodes in the tree, just gives rise to incorrect values, even if Bellman error is fully minimized at those states. Thus, most of the Bellman updates do not actually bring Q-values at the states of the MDP closer to <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-285dd7f4fb9e9c4e2f9bc31308771a4e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#94;&#42;" title="Rendered by QuickLaTeX.com" height="16" width="20" style="vertical-align: -4px;"/>, since the primary cause of incorrect target values isn’t corrected.</p>
<p>This observation is surprising, since it demonstrates how the choice of an online distribution coupled with function approximation might actually learn incorrect Q-values. On the other hand, a scheme that chooses to update states level by level progressively (Figure 2), ensuring that target values used at any iteration of learning are correct, very easily learns correct Q-values in this example.</p>
<p><!--
![Figure 2: Run of an ADP algorithm with an oracle distribution, that updates states level-by level, progressing through the tree from the leaves to the root. Even in the presence of function approximation, selecting the right set of nodes for updates gives rise to correct Q-values.](https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1584341988034_discor_func_approx_final.png)
--></p>
<p style="text-align:center;">
<img decoding="async" src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1584341988034_discor_func_approx_final.png" width="" /><br />
<br />
<i><br />
Figure 2:  Run of an ADP algorithm with an oracle distribution, that updates states level-by level, progressing through the tree from the leaves to the root. Even in the presence of function approximation, selecting the right set of nodes for updates gives rise to correct Q-values.<br />
</i>
</p>
<h1 id="consequences-of-absent-corrective-feedback">Consequences of Absent Corrective Feedback</h1>
<p>Now, one might ask if an absence of corrective feedback occurs in practice, beyond a simple didactic example and whether it hurts in practical problems. Since visualizing the dynamics of the learning process is hard in practical problems as we did for the didactic example, we instead devise a metric that quantifies our intuition for corrective feedback. This metric, what we call <em>value error,</em> is given by:</p>
<p><script type="math/tex; mode=display">\mathcal{E}_k = \mathbb{E}_{d^{\pi_k}} [|Q_k - Q^*|]</script></p>
<p>Increasing values of <script type="math/tex">\mathcal{E}_k</script> imply that the algorithm is pushing Q-values farther away from <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-285dd7f4fb9e9c4e2f9bc31308771a4e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#94;&#42;" title="Rendered by QuickLaTeX.com" height="16" width="20" style="vertical-align: -4px;"/>, which means that corrective feedback is absent, if this happens over a number of iterations. On the other hand, decreasing values of <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-70b088a98ba47084cea40b0dfea50768_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#69;&#125;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="15" width="16" style="vertical-align: -3px;"/> implies that the algorithm is continuously improving its estimate of <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-2c758bec4c272382411b95fc0e7ee250_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;" title="Rendered by QuickLaTeX.com" height="16" width="14" style="vertical-align: -4px;"/>, by moving it towards <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-285dd7f4fb9e9c4e2f9bc31308771a4e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#81;&#94;&#42;" title="Rendered by QuickLaTeX.com" height="16" width="20" style="vertical-align: -4px;"/> with each iteration, indicating the presence of corrective feedback.</p>
<p>Observe in Figure 3, that ADP methods can suffer from prolonged periods where this global measure of error in the Q-function, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-70b088a98ba47084cea40b0dfea50768_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#69;&#125;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="15" width="16" style="vertical-align: -3px;"/>, is increasing or fluctuating, and the corresponding returns degrade or stagnate, implying an absence of corrective feedback.</p>
<p><!--
![Figure 3: Consequences of absent corrective feedback, including (a) sub-optimal convergence, (b) instability in learning and (c) inability to learn with sparse rewards.](https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1583775913964_Screen+Shot+2020-03-09+at+10.44.58+AM.png)
--></p>
<p style="text-align:center;">
<img decoding="async" src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1583775913964_Screen+Shot+2020-03-09+at+10.44.58+AM.png" width="" /><br />
<br />
<i><br />
Figure 3: Consequences of absent corrective feedback, including (a) sub-optimal convergence, (b) instability in learning and (c) inability to learn with sparse rewards.<br />
</i>
</p>
<p>In particular, we describe three different consequences of an absence of corrective feedback:</p>
<ol>
<li>
<p><strong>Convergence to suboptimal Q-functions.</strong> We find that on-policy sampling can cause ADP to converge to a suboptimal solution, even in the absence of sampling error. Figure 3(a) shows that the value error <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-422a9040d625862c3f8d14b29b45a8c5_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#69;&#125;&#95;&#107;&#60;" title="Rendered by QuickLaTeX.com" height="15" width="36" style="vertical-align: -3px;"/> rapidly decreases initially, and eventually converges to a value significantly greater than 0, from which the learning process never recovers.</p>
</li>
<li>
<p><strong>Instability in the learning process.</strong> We observe that ADP with replay buffers can be unstable. For instance, the algorithm is prone to degradation even if the latest policy obtains returns that are very close to the optimal return in Figure 3(b).</p>
</li>
<li>
<p><strong>Inability to learn with low signal-to-noise ratio.</strong> Absence of corrective feedback can also prevent ADP algorithms from learning quickly in scenarios with low signal-to-noise ratio, such as tasks with sparse/noisy rewards as shown in Figure 3(c). Note that this is not an exploration issue, since all transitions in the MDP are provided to the algorithm in this experiment.</p>
</li>
</ol>
<h1 id="inducing-maximal-corrective-feedback-via-distribution-correction">Inducing Maximal Corrective Feedback via Distribution Correction</h1>
<p>Now that we have defined corrective feedback and gone over some detrimental consequences an absence of it can have on the learning process of an ADP algorithm, what might be some ways to fix this problem? To recap, an absence of corrective feedback occurs when ADP algorithms naively use the on-policy or replay buffer distributions for training Q-functions. One way to prevent this problem is by computing an “optimal” data distribution that provides maximal corrective feedback, and train Q-functions using this distribution? This way we can ensure that the ADP algorithm always enjoys corrective feedback, and hence makes steady learning progress. The strategy we used in our work is to compute this optimal distribution and then perform a <strong>weighted Bellman update</strong> that re-weights the data distribution in the replay buffer to this optimal distribution (in practice, a tractable approximation is required, as we will see) via importance sampling based techniques.</p>
<p>We will not go into the full details of our derivation in this article, however, we mention the optimization problem used to obtain a form for this optimal distribution and encourage readers interested in the theory to checkout Section 4 in our paper.  In this optimization problem, our goal is to minimize a measure of corrective feedback, given by <em>value</em> error <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-70b088a98ba47084cea40b0dfea50768_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#69;&#125;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="15" width="16" style="vertical-align: -3px;"/>, with respect to the distribution <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-246a87e28e6d0114e2442071efcab646_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#112;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="12" width="17" style="vertical-align: -4px;"/> used for Bellman error minimization, at every iteration <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3422b6bb5c160593658b7c39425d9880_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#107;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: 0px;"/>. This gives rise to the following problem:</p>
<p><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-6612c05cda14eb379821ce1b9993ecd3_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#109;&#105;&#110;&#32;&#95;&#123;&#112;&#95;&#123;&#107;&#125;&#125;&#32;&#92;&#59;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#69;&#125;&#95;&#123;&#100;&#94;&#123;&#92;&#112;&#105;&#95;&#123;&#107;&#125;&#125;&#125;&#91;&#124;&#81;&#95;&#123;&#107;&#125;&#45;&#81;&#94;&#123;&#42;&#125;&#124;&#93;" title="Rendered by QuickLaTeX.com" height="20" width="171" style="vertical-align: -6px;"/></p>
<p><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-a30e53297d15d1e2167868ef267b6dea_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#116;&#101;&#120;&#116;&#32;&#123;&#32;&#115;&#46;&#116;&#46;&#32;&#125;&#92;&#59;&#92;&#59;&#32;&#32;&#81;&#95;&#123;&#107;&#125;&#61;&#92;&#97;&#114;&#103;&#32;&#92;&#109;&#105;&#110;&#32;&#95;&#123;&#81;&#125;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#69;&#125;&#95;&#123;&#112;&#95;&#123;&#107;&#125;&#125;&#92;&#108;&#101;&#102;&#116;&#91;&#92;&#108;&#101;&#102;&#116;&#40;&#81;&#45;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#66;&#125;&#94;&#123;&#42;&#125;&#32;&#81;&#95;&#123;&#107;&#45;&#49;&#125;&#92;&#114;&#105;&#103;&#104;&#116;&#41;&#94;&#123;&#50;&#125;&#92;&#114;&#105;&#103;&#104;&#116;&#93;" title="Rendered by QuickLaTeX.com" height="32" width="320" style="vertical-align: -11px;"/></p>
<p>We show in our paper that the solution of this optimization problem, that we refer to as the optimal distribution, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-9745dad6b7f279e5bf4a6045defa7bac_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#112;&#95;&#107;&#94;&#42;" title="Rendered by QuickLaTeX.com" height="17" width="17" style="vertical-align: -5px;"/>, is given by:</p>
<p><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3e7f10c4137f316f056dc83105f2ca98_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#112;&#95;&#123;&#107;&#125;&#94;&#42;&#40;&#115;&#44;&#32;&#97;&#41;&#32;&#92;&#112;&#114;&#111;&#112;&#116;&#111;&#32;&#92;&#101;&#120;&#112;&#32;&#92;&#108;&#101;&#102;&#116;&#40;&#45;&#92;&#108;&#101;&#102;&#116;&#124;&#81;&#95;&#123;&#107;&#125;&#45;&#81;&#94;&#123;&#42;&#125;&#92;&#114;&#105;&#103;&#104;&#116;&#124;&#40;&#115;&#44;&#32;&#97;&#41;&#92;&#114;&#105;&#103;&#104;&#116;&#41;&#32;&#92;&#102;&#114;&#97;&#99;&#123;&#92;&#108;&#101;&#102;&#116;&#124;&#81;&#95;&#123;&#107;&#125;&#45;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#66;&#125;&#94;&#123;&#42;&#125;&#32;&#81;&#95;&#123;&#107;&#45;&#49;&#125;&#92;&#114;&#105;&#103;&#104;&#116;&#124;&#40;&#115;&#44;&#32;&#97;&#41;&#125;&#123;&#92;&#108;&#97;&#109;&#98;&#100;&#97;&#94;&#123;&#42;&#125;&#125;" title="Rendered by QuickLaTeX.com" height="25" width="380" style="vertical-align: -6px;"/></p>
<p>By simplifying this expression, we obtain a practically viable expression for weights, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-1df460785d846f8cf32198526f5a271e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#119;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="11" width="20" style="vertical-align: -3px;"/>, at any iteration <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3422b6bb5c160593658b7c39425d9880_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#107;" title="Rendered by QuickLaTeX.com" height="12" width="9" style="vertical-align: 0px;"/> that can be used to re-weight the data distribution:</p>
<p><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3938760a543e41788a53be8bba9bcc46_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#119;&#95;&#123;&#107;&#125;&#40;&#115;&#44;&#32;&#97;&#41;&#32;&#92;&#112;&#114;&#111;&#112;&#116;&#111;&#32;&#92;&#101;&#120;&#112;&#32;&#92;&#108;&#101;&#102;&#116;&#40;&#45;&#92;&#102;&#114;&#97;&#99;&#123;&#92;&#103;&#97;&#109;&#109;&#97;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#69;&#125;&#95;&#123;&#115;&#39;&#124;&#115;&#44;&#32;&#97;&#125;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#69;&#125;&#95;&#123;&#97;&#39;&#32;&#92;&#115;&#105;&#109;&#32;&#92;&#112;&#105;&#95;&#92;&#112;&#104;&#105;&#40;&#92;&#99;&#100;&#111;&#116;&#124;&#115;&#39;&#41;&#125;&#32;&#92;&#68;&#101;&#108;&#116;&#97;&#95;&#123;&#107;&#45;&#49;&#125;&#40;&#115;&#39;&#44;&#32;&#97;&#39;&#41;&#125;&#123;&#92;&#116;&#97;&#117;&#125;&#92;&#114;&#105;&#103;&#104;&#116;&#41;" title="Rendered by QuickLaTeX.com" height="43" width="344" style="vertical-align: -17px;"/></p>
<p>where <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3a4e60efde2d18bf977915da8b614b7f_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#68;&#101;&#108;&#116;&#97;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="16" width="22" style="vertical-align: -3px;"/> is the accumulated Bellman error over iterations, and it satisfies a convenient recursion making it amenable to practical implementations,</p>
<p><img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-adaab580581cb5b638ec0b60e5d221cb_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#68;&#101;&#108;&#116;&#97;&#95;&#123;&#107;&#125;&#40;&#115;&#44;&#32;&#97;&#41;&#32;&#61;&#92;&#108;&#101;&#102;&#116;&#124;&#81;&#95;&#123;&#107;&#125;&#45;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#66;&#125;&#94;&#123;&#42;&#125;&#32;&#81;&#95;&#123;&#107;&#45;&#49;&#125;&#92;&#114;&#105;&#103;&#104;&#116;&#124;&#40;&#115;&#44;&#32;&#97;&#41;&#32;&#43;&#92;&#103;&#97;&#109;&#109;&#97;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#69;&#125;&#95;&#123;&#115;&#39;&#124;&#115;&#44;&#32;&#97;&#125;&#32;&#92;&#109;&#97;&#116;&#104;&#98;&#98;&#123;&#69;&#125;&#95;&#123;&#97;&#39;&#32;&#92;&#115;&#105;&#109;&#32;&#92;&#112;&#105;&#95;&#92;&#112;&#104;&#105;&#40;&#92;&#99;&#100;&#111;&#116;&#124;&#115;&#39;&#41;&#125;&#32;&#92;&#68;&#101;&#108;&#116;&#97;&#95;&#123;&#107;&#45;&#49;&#125;&#40;&#115;&#39;&#44;&#32;&#97;&#39;&#41;" title="Rendered by QuickLaTeX.com" height="22" width="487" style="vertical-align: -8px;"/></p>
<p>and <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-aa37ac11527640d479cf10422b69c87e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#112;&#105;&#95;&#92;&#112;&#104;&#105;" title="Rendered by QuickLaTeX.com" height="14" width="18" style="vertical-align: -6px;"/> is the Boltzmann or <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-ef559fe07ca6620424abd6f3c9444bde_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#118;&#97;&#114;&#101;&#112;&#115;&#105;&#108;&#111;&#110;&#45;" title="Rendered by QuickLaTeX.com" height="8" width="21" style="vertical-align: 0px;"/>greedy policy corresponding to the current Q function.</p>
<p>What does this expression for <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-1df460785d846f8cf32198526f5a271e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#119;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="11" width="20" style="vertical-align: -3px;"/> intuitively correspond to? Observe that the term appearing in the exponent in the expression for <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-1df460785d846f8cf32198526f5a271e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#119;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="11" width="20" style="vertical-align: -3px;"/> corresponds to the accumulated Bellman error in the target values. Our choice of <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-1df460785d846f8cf32198526f5a271e_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#119;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="11" width="20" style="vertical-align: -3px;"/>, thus, basically down-weights transitions with highly incorrect target values. This technique falls into a broader class of <strong>abstention</strong> based techniques that are common in supervised learning settings with noisy labels, where down-weighting datapoints (transitions here) with errorful labels (target values here) can boost generalization and correctness properties of the learned model.</p>
<p><!--
![Figure 4: Schematic of the DisCor algorithm. Transitions with errorful target values are downweighted.](https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1584342258383_discor_schematic.png)
--></p>
<p style="text-align:center;">
<img decoding="async" src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1584342258383_discor_schematic.png" width="" /><br />
<br />
<i><br />
Figure 4: Schematic of the DisCor algorithm. Transitions with errorful target values are downweighted.<br />
</i>
</p>
<p>Why does our choice of <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3a4e60efde2d18bf977915da8b614b7f_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#68;&#101;&#108;&#116;&#97;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="16" width="22" style="vertical-align: -3px;"/>, i.e. the sum of accumulated Bellman errors suffice? This is because this value <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3a4e60efde2d18bf977915da8b614b7f_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#68;&#101;&#108;&#116;&#97;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="16" width="22" style="vertical-align: -3px;"/> accounts for how error is propagated in ADP methods. Bellman errors, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-b9b09c72379d2340aa806a89813e0b97_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#124;&#81;&#95;&#107;&#32;&#45;&#32;&#92;&#109;&#97;&#116;&#104;&#99;&#97;&#108;&#123;&#66;&#125;&#94;&#42;&#81;&#95;&#123;&#107;&#45;&#49;&#125;&#124;" title="Rendered by QuickLaTeX.com" height="19" width="110" style="vertical-align: -5px;"/> are propagated under the current policy <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-747ea2dfedeab4a4e158ea8f47023043_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#112;&#105;&#95;&#123;&#107;&#45;&#49;&#125;" title="Rendered by QuickLaTeX.com" height="11" width="35" style="vertical-align: -3px;"/>, and then discounted when computing target values for updates in ADP. <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-3a4e60efde2d18bf977915da8b614b7f_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#92;&#68;&#101;&#108;&#116;&#97;&#95;&#107;" title="Rendered by QuickLaTeX.com" height="16" width="22" style="vertical-align: -3px;"/> captures exactly this, and therefore, using this estimate in our weights suffices.</p>
<p>Our practical algorithm, that we refer to as <strong>D</strong>is<strong>C</strong>or (Distribution Correction), is identical to conventional ADP methods like Q-learning, with the exception that it performs a weighted Bellman backup – it assigns a weight <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-18d29b564c6c28e47c7dcbca83372089_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#119;&#95;&#107;&#40;&#115;&#44;&#97;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="59" style="vertical-align: -5px;"/> to a transition, <img decoding="async" src="https://robohub.org/wp-content/ql-cache/quicklatex.com-6a9a3ccb9f256cd97473514276438e95_l3.png" class="ql-img-inline-formula quicklatex-auto-format" alt="&#40;&#115;&#44;&#32;&#97;&#44;&#32;&#114;&#44;&#32;&#115;&#39;&#41;" title="Rendered by QuickLaTeX.com" height="19" width="74" style="vertical-align: -5px;"/> and performs a Bellman backup weighted by these weights, as shown below.</p>
<p><script type="math/tex; mode=display">Q_k \leftarrow \arg \min_Q \frac{1}{N} \sum_{i=1}^N w_i(s, a) \cdot \left(Q(s, a) - [r(s, a) + \gamma Q_{k-1}(s', a')]\right)^2</script></p>
<p>We depict the general principle in the schematic diagram shown in Figure 4.</p>
<h1 id="how-does-discor-perform-in-practice">How does DisCor perform in practice?</h1>
<p>We finally present some results that demonstrate the efficacy of our method, DisCor, in practical scenarios. Since DisCor only modifies the chosen distribution for the Bellman update, it can be applied on top of any standard ADP algorithm including soft actor-critic (SAC) or deep Q-network (DQN). Our paper presents results for a number of tasks spanning a wide variety of settings including robotic manipulation tasks, multi-task reinforcement learning tasks, learning with stochastic and noisy rewards, and Atari games. In this blog post, we present two of these results from robotic manipulation and multi-task RL.</p>
<ol>
<li>
<p><strong>Robotic manipulation tasks.</strong> On six challenging benchmark tasks from the <a href="https://meta-world.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">MetaWorld</a> suite, we observe that DisCor when combined with SAC greatly outperforms prior state-of-the-art RL algorithms such as soft actor-critic (SAC) and prioritized experience replay (<a href="https://arxiv.org/abs/1511.05952" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PER</a>) which is a prior method that prioritizes states with high Bellman error during training. Note that DisCor usually starts learning earlier than other methods compared to. DisCor outperforms vanilla SAC by a factor of about <strong>50%</strong> on average, in terms of success rate on these tasks.</p>
<p style="text-align:center;">
<img decoding="async" src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1583776003015_Screen+Shot+2020-03-09+at+10.45.54+AM.png" width="" /><br />

</p>
</li>
<li>
<p><strong>Multi-task reinforcement learning.</strong> We also present certain results on the Multi-task 10 (MT10) and Multi-task 50 (MT50) benchmarks from the Meta-world suite. The goal here is to learn a single policy that can solve a number of (10 or 50, respectively) different manipulation tasks that share common structure.  We note that DisCor outperforms, state-of-the-art SAC algorithm on both of these benchmarks by a wide margin (for e.g. <strong>50%</strong> on MT10, success rate). Unlike the learning process of SAC that tends to plateau over the course of learning, we observe that DisCor always exhibits a non-zero gradient for the learning process, until it converges.</p>
<p style="text-align:center;">
<img decoding="async" src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1584342410332_Screen+Shot+2020-03-16+at+12.06.22+AM.png" width="" /><br />

</p>
</li>
</ol>
<p><!--
*
--></p>
<p>In our paper, we also perform evaluations on other domains such as Atari games and OpenAI gym benchmarks, and we encourage the readers to check those out. We also perform an analysis of the method on tabular domains, understanding different aspects of the method.</p>
<h1 id="perspectives-future-work-and-open-problems">Perspectives, Future Work and Open Problems</h1>
<p>Some of <a href="https://arxiv.org/abs/1906.00949" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our</a> and <a href="https://arxiv.org/abs/1712.06924" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">other</a> prior work has highlighted the impact of the data distribution on the performance of ADP algorithms, We observed in another <a href="https://arxiv.org/abs/1902.10250" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">prior work</a> that in contrast to the intuitive belief about the efficacy of online Q-learning with on-policy data collection, Q-learning with a uniform distribution over states and actions seemed to perform best. Obtaining a uniform distribution over state-action tuples during training is not possible in RL, unless all states and actions are observed at least once, which may not be the case in a number of scenarios. We might also ask the question about whether the uniform distribution is the best choice that can be used in an RL setting? The form of the optimal distribution derived in Section 4 of our paper, is a potentially better choice since it is customized to the MDP under consideration.</p>
<p>Furthermore, in the domain of purely offline reinforcement learning, studied in <a href="https://arxiv.org/abs/1906.00949" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our</a> prior work and some other works, such as <a href="https://arxiv.org/abs/1712.06924" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this</a> and <a href="https://arxiv.org/abs/1907.04543" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this</a>, we observe that the data distribution is again a central feature, where backing up out-of-distribution actions and the inability to try these actions out in the environment to obtain answers to counterfactual queries, can cause error accumulation and backups to diverge. However, in this work, we demonstrate a somewhat counterintuitive finding: even with on-policy data collection, where the algorithm, in principle, can evaluate all forms of counterfactual queries, the algorithm may not obtain a steady learning progress, due to an undesirable interaction between the data distribution and generalization effects of the function approximator.</p>
<p style="text-align:center;">
<img decoding="async" src="https://paper-attachments.dropbox.com/s_B1B66C162F9CE3C53CD9315CF7237DAB864CB3AAEED950E42789F193B67C60EB_1583852356782_batchgraph.png" width="" /><br />

</p>
<h2 id="what-might-be-a-few-promising-directions-to-pursue-in-future-work">What might be a few promising directions to pursue in future work?</h2>
<p><strong>Formal analysis of learning dynamics:</strong> While our study is an initial foray into the role that data distributions play in the learning dynamics of ADP algorithms, this motivates a significantly deeper direction of future study. We need to answer questions related to how deep neural network based function approximators actually behave, which are behind these ADP methods, in order to get them to enjoy corrective feedback.</p>
<p><strong>Re-weighting to supplement exploration in RL problems:</strong> Our work depicts the promise of re-weighting techniques as a practically simple replacement for altering entire exploration strategies. We believe that re-weighting techniques are very promising as a general tool in our toolkit to develop RL algorithms. In an online RL setting, re-weighting can help remove the some of the burden off exploration algorithms, and can thus, potentially help us employ complex exploration strategies in RL algorithms.</p>
<p>More generally, we would like to make a case of analyzing effects of data distribution more deeply in the context of deep RL algorithms. It is well known that narrow distributions can lead to brittle solutions in supervised learning that also do not generalize. What is the corresponding analogue in reinforcement learning? Distributional robustness style techniques have been used in supervised learning to guarantee a uniformly convergent learning process, but it still remains unclear how to apply these in an RL with function approximation setting. Part of the reason is that the theory of RL often derives from tabular settings, where distributions do not hamper the learning process to the extent they do with function approximation. However, as we showed in this work, choosing the right distribution may lead to significant gains in deep RL methods, and therefore, we believe, that this issue should be studied in more detail.</p>
<hr />
<p>This blog post is based on our recent paper:</p>
<ul>
<li><strong>DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction</strong> <br />
Aviral Kumar, Abhishek Gupta, Sergey Levine <br />
<a href="https://arxiv.org/abs/2003.07305" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">arXiv</a></li>
</ul>
<p>We thank Sergey Levine and Marvin Zhang for their valuable feedback on this blog post.</p>
<hr />
<p>This article was initially published on the <a href="https://bair.berkeley.edu/blog/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Emergent behavior by minimizing chaos</title>
		<link>https://robohub.org/emergent-behavior-by-minimizing-chaos/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sun, 26 Jan 2020 23:13:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/emergent-behavior-by-minimizing-chaos/</guid>

					<description><![CDATA[


All living organisms carve out environmental niches within which they can
maintain relative predictability amidst the ever-increasing entropy around them
(1),
(2).
Humans, for example, go to great lengths to shield themselves from surprise —
we b...]]></description>
										<content:encoded><![CDATA[<p><strong>By Glen Berseth</strong><br />
<meta name="twitter:title" content="SMiRL: Surprise Minimizing RL in Dynamic Environments" />  <meta name="twitter:card" content="summary_image" />  <meta name="twitter:image" content="https://bair.berkeley.edu/static/blog/smirl/SMiRL_Outline.png" />  </p>
<p>All living organisms carve out environmental niches within which they can maintain relative predictability amidst the ever-increasing entropy around them <a href="http://www.ler.esalq.usp.br/aulas/lce1302/life_as_a_manifestation.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">(1)</a>, <a href="https://www.fil.ion.ucl.ac.uk/~karl/The%20free-energy%20principle%20-%20a%20rough%20guide%20to%20the%20brain.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">(2)</a>. Humans, for example, go to great lengths to shield themselves from surprise — we band together in millions to build cities with homes, supplying water, food, gas, and electricity to control the deterioration of our bodies and living spaces amidst heat and cold, wind and storm. The need to discover and maintain such surprise-free equilibria has driven great resourcefulness and skill in organisms across very diverse natural habitats. Motivated by this, we ask: could the motive of preserving order amidst chaos guide the automatic acquisition of useful behaviors in artificial agents?</p>
<p>  <span id="more-157317"></span>  </p>
<p>How might an agent in an environment acquire complex behaviors and skills with no external supervision? This central problem in artificial intelligence has evoked several candidate solutions, largely focusing on novelty-seeking behaviors <a href="http://people.idsia.ch/~juergen/curioussingapore/curioussingapore.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">(3)</a>, <a href="https://arxiv.org/abs/1606.01868" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">(4)</a>, <a href="https://pathak22.github.io/noreward-rl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">(5)</a>. In simulated worlds, such as video games, novelty-seeking intrinsic motivation can lead to interesting and meaningful behavior. However, these environments may be fundamentally lacking compared to the real world. In the real world, natural forces and other agents offer bountiful novelty. Instead, the challenge in natural environments is allostasis: discovering behaviors that enable agents to maintain an equilibrium (homeostasis), for example to preserve their bodies, their homes, and avoid predators and hunger. In the example below we shown an example where an agent is experiencing random events due to the changing weather. If the agent learns to build a shelter, in this case a house, the agent will reduce the observed effects from weather.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/smirl/robotsurprise_stacked.png" width="600" />  </p>
<p>We formalize homeostasis as an objective for reinforcement learning based on surprise minimization (SMiRL). In entropic and dynamic environments with undesirable forms of novelty, minimizing surprise (i.e., minimizing novelty) causes agents to naturally seek an equilibrium that can be stably maintained.</p>
<p style="text-align:center;"> <img decoding="async" width="600" src="https://bair.berkeley.edu/static/blog/smirl/SMiRL_Outline.png" />  </p>
<p>Here we show an illustration of the agent interaction loop using SMiRL. When the agent observes a state $\mathbf{s}$, it computes the probability of this new state given the belief the agent has $r_{t} \leftarrow p_{\theta_{t-1}}(\textbf{s})$. This belief models the states the agent is most familiar with – i.e., the distribution of states it has seen in the past. Experiencing states that are more familiar will result in higher reward. After the agent experience a new state it updates its belief $p_{\theta_{t-1}}(\textbf{s})$ over states to include the most recent experience. Then, the goal of the action policy $\pi(a|\textbf{s}, \theta_{t})$ is to choose actions that will result in the agent consistently experiencing familiar states. Crucially, the agent understands that its beliefs will change in the future. This means that it has two mechanisms by which to maximize this reward: taking actions to visit familiar states, and taking actions to visit states that will <em>change its beliefs</em> such that future states are more familiar. It is this latter mechanism that results in complex emergent behavior. Below, we visualize a policy trained to play the game of Tetris. On the left the blocks the agent chooses are shown and on the right is a visualization of $p_{\theta_{t}}(\textbf{s})$. We can see how as the episode progresses the belief over possible locations to place blocks tends to favor only the bottom row. This encourages the agent to eliminate blocks to prevent board from filling up.</p>
<p>  <!-- 

<div class="t">     

<table align="center">         

<tr>     

<td align="center">         <img decoding="async" width="200" src="https://bair.berkeley.edu/static/blog/smirl/tetris_ps.gif">         </td>

     

<td>     <img decoding="async" width="320" src="https://bair.berkeley.edu/static/blog/smirl/minigrid-maze-random-count.gif">            </td>

 	</tr>

         

<tr align=center>         

<td>             Tetris             </td>

         

<td>             HauntedHouse             </td>

         </tr>

 </table>

 </div>

 -->  </p>
<p style="text-align:center;"> <img decoding="async" height="200" src="https://bair.berkeley.edu/static/blog/smirl/tetris_ps.gif" /> <img decoding="async" height="220" src="https://bair.berkeley.edu/static/blog/smirl/minigrid-maze-random-count.gif" /> <br /> <i> Left: Tetris. Right: HauntedHouse. </i> </p>
<h4 id="emergent-behavior">Emergent behavior</h4>
<p>The SMiRL agent demonstrates meaningful emergent behaviors in a number of different environments. In the Tetris environment, the agent is able to learn proactive behaviors to eliminate rows and properly play the game. The agent also learns emergent game playing behavior in the VizDoom environment, acquiring an effective policy for dodging the fireballs thrown by the enemies. In both of these environments, stochastic and chaotic events force the SMiRL agent to take a coordinated course of action to avoid unusual states, such as full Tetris boards or fireball explorations.</p>
<p>  <!-- | Doom Hold The Line                                           | Doom Defend The Line                                         |                         HauntedHouse                         | | ------------------------------------------------------------ | ------------------------------------------------------------ | :----------------------------------------------------------: | | <img decoding="async" width="100%" src="https://bair.berkeley.edu/static/blog/smirl/Doom_trained_enough_result.gif"> | <img decoding="async" width="100%" src="https://bair.berkeley.edu/static/blog/smirl/vizdoom_dtl.gif"> | <img decoding="async" width="70%" src="https://bair.berkeley.edu/static/blog/smirl/minigrid-maze-random-count.gif"> |  

<div class="t">     

<table align="center">         

<tr>     

<td>     <img decoding="async" width="100%" src="https://bair.berkeley.edu/static/blog/smirl/Doom_trained_enough_result.gif">         </td>

     

<td>     <img decoding="async" width="100%" src="https://bair.berkeley.edu/static/blog/smirl/vizdoom_dtl.gif">            </td>

 	</tr>

         

<tr align=center>         

<td>             Doom Hold The Line             </td>

         

<td>             Doom Defend The Line             </td>

         </tr>

 </table>

 </div>

 -->  </p>
<p style="text-align:center;"> <img decoding="async" width="45%" src="https://bair.berkeley.edu/static/blog/smirl/Doom_trained_enough_result.gif" /> <img decoding="async" width="45%" src="https://bair.berkeley.edu/static/blog/smirl/vizdoom_dtl.gif" /> <br /> <i> Left: Doom Hold The Line. Right: Doom Defend The Line. </i> </p>
<h5 id="biped">Biped</h5>
<p>In the Cliff environment, the agent learns a policy that greatly reduces the probability of falling off of the cliff by bracing against the ground and stabilize itself at the edge, as shown in the figure below. In the <em>Treadmill</em> environment, SMiRL learns a more complex locomotion behavior, jumping forward to increase the time it stays on the treadmill, as shown in figure below.</p>
<p>  <!-- 

<div class="t">     

<table align="center">         

<tr>     

<td>         <video width="320" height="240" autoplay><source src="https://bair.berkeley.edu/static/blog/smirl/cliff_surpise_VAE_6_v3_rewardViz.mp4" type="video/mp4"><source src="movie.ogg" type="video/ogg">Your browser does not support the video tag.</video>         </td>

     

<td>     <video width="320" height="240" autoplay><source src="https://bair.berkeley.edu/static/blog/smirl/treadmill_surpise_VAE_6_v3_rewardViz.mp4" type="video/mp4"><source src="movie.ogg" type="video/ogg">Your browser does not support the video tag.</video>            </td>

 	</tr>

         

<tr align=center>         

<td>             Cliff             </td>

         

<td>             Treadmill             </td>

         </tr>

 </table>

 </div>

  ​ -->  </p>
<p style="text-align:center;"> <video width="320" height="240" style="margin: 10px;" autoplay=""><source src="https://bair.berkeley.edu/static/blog/smirl/cliff_surpise_VAE_6_v3_rewardViz.mp4" type="video/mp4" /><source src="movie.ogg" type="video/ogg" />Your browser does not support the video tag.</video> <video width="320" height="240" style="margin: 10px;" autoplay=""><source src="https://bair.berkeley.edu/static/blog/smirl/treadmill_surpise_VAE_6_v3_rewardViz.mp4" type="video/mp4" /><source src="movie.ogg" type="video/ogg" />Your browser does not support the video tag.</video> <br /> <i> Left: Cliff. Right: Treadmill. </i> </p>
<h4 id="comparison-to-intrinsic-motivation">Comparison to Intrinsic motivation:</h4>
<p>Intrinsic motivation is the idea that behavior is driven by internal reward signals that are task independent. Below, we show plots of the environment-specific rewards over time on Tetris, VizDoomTakeCover, and the humanoid domains. In order to compare SMiRL to more standard intrinsic motivation methods, which seek out states that maximize surprise or novelty, we also evaluated ICM <a href="https://pathak22.github.io/noreward-rl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">(5)</a> and RND <a href="https://arxiv.org/abs/1810.12894" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">(6)</a>. We include an oracle agent that directly optimizes the task reward. On Tetris, after training for $2000$ epochs, SMiRL achieves near perfect play, on par with the oracle reward optimizing agent, with no deaths. ICM seeks novelty by creating more and more distinct patterns of blocks rather than clearing them, leading to deteriorating game scores over time. On VizDoomTakeCover, SmiRL effectively learns to dodge fireballs thrown by the adversaries.</p>
<p style="text-align:center;"> <img decoding="async" width="90%" src="https://bair.berkeley.edu/static/blog/smirl/video_game_comparisons_2.png" />  </p>
<p>The baseline comparisons for the Cliff and Treadmill environments have a similar outcome. The novelty seeking behavior of ICM causes it to learn a type of irregular behavior that causes the agent to jump off the Cliff and roll around on the Treadmill, maximizing the variety (and quantity) of falls.</p>
<h4 id="smirl--curiosity">SMiRL + Curiosity:</h4>
<p style="text-align:center;"> <img decoding="async" width="90%" src="https://bair.berkeley.edu/static/blog/smirl/Capture_biped_results.png" />  </p>
<p>While on the surface, SMiRL minimizes surprise and curiosity approaches like ICM maximize novelty, they are in fact not mutually incompatible. In particular, while ICM maximizes novelty with respect to a learned transition model, SMiRL minimizes surprise with respect to a learned state distribution. We can combine ICM and SMiRL to achieve even better results on the Treadmill environment.</p>
<p>  <!-- 

<div class="t">     

<table align="center">         

<tr>     

<td>         <video width="320" height="240" autoplay><source src="https://bair.berkeley.edu/static/blog/smirl/treadmill_surpise_ICM_v3_rewardViz.mp4" type="video/mp4"><source src="movie.ogg" type="video/ogg">Your browser does not support the video tag.</video>         </td>

     

<td>     <video width="320" height="240" autoplay><source src="https://bair.berkeley.edu/static/blog/smirl/pedistal_surpise_v3_rewardViz.mp4" type="video/mp4"><source src="movie.ogg" type="video/ogg">Your browser does not support the video tag.</video>            </td>

 	</tr>

         

<tr align=center>         

<td>             Treadmill + ICM             </td>

         

<td>             Pedestal             </td>

         </tr>

 </table>

 </div>

 -->  </p>
<p style="text-align:center;"> <video width="320" height="240" style="margin: 10px;" autoplay=""><source src="https://bair.berkeley.edu/static/blog/smirl/treadmill_surpise_ICM_v3_rewardViz.mp4" type="video/mp4" /><source src="movie.ogg" type="video/ogg" />Your browser does not support the video tag.</video> <video width="320" height="240" style="margin: 10px;" autoplay=""><source src="https://bair.berkeley.edu/static/blog/smirl/pedistal_surpise_v3_rewardViz.mp4" type="video/mp4" /><source src="movie.ogg" type="video/ogg" />Your browser does not support the video tag.</video> <br /> <i> Left: Treadmill+ICM. Right: Pedestal. </i> </p>
<p>  <!-- 

<div class="containerWide">   

<div class="photosWide">     <video width="320" height="240" style="margin: 20px;" autoplay><source src="https://bair.berkeley.edu/static/blog/smirl/treadmill_surpise_ICM_v3_rewardViz.mp4" type="video/mp4"><source src="movie.ogg" type="video/ogg">Your browser does not support the video tag.</video>     <span class="wordWide"><i>Treadmill+ICM.</i></span>   </div>

    

<div class="photosWide">     <video width="320" height="240" style="margin: 20px;" autoplay><source src="https://bair.berkeley.edu/static/blog/smirl/pedistal_surpise_v3_rewardViz.mp4" type="video/mp4"><source src="movie.ogg" type="video/ogg">Your browser does not support the video tag.</video>       <span class="wordWide"><i>Pedestal</i></span>   </div>

 </div>

 -->  </p>
<h4 id="insights">Insights:</h4>
<p>The key insight utilized by our method is that, in contrast to simple simulated domains, realistic environments exhibit dynamic phenomena that gradually increase entropy over time. An agent that resists this growth in entropy must take active and coordinated actions, thus learning increasingly complex behaviors. This is different from commonly proposed intrinsic exploration methods based on novelty, which instead seek to visit novel states and increase entropy. SMiRL holds promise for a new kind of unsupervised RL method that produces behaviors that are closely tied to the prevailing disruptive forces, adversaries, and other sources of entropy in the environment.</p>
<p>This article was initially published on the <a href="https://bair.berkeley.edu/blog/2019/12/18/smirl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the author’s permission.</p>
]]></content:encoded>
					
		
		<enclosure url="https://bair.berkeley.edu/static/blog/smirl/cliff_surpise_VAE_6_v3_rewardViz.mp4" length="1085487" type="video/mp4" />
<enclosure url="https://bair.berkeley.edu/static/blog/smirl/treadmill_surpise_VAE_6_v3_rewardViz.mp4" length="3141454" type="video/mp4" />
<enclosure url="https://bair.berkeley.edu/static/blog/smirl/treadmill_surpise_ICM_v3_rewardViz.mp4" length="5746898" type="video/mp4" />
<enclosure url="https://bair.berkeley.edu/static/blog/smirl/pedistal_surpise_v3_rewardViz.mp4" length="4342590" type="video/mp4" />

			</item>
		<item>
		<title>Data-driven deep reinforcement learning</title>
		<link>https://robohub.org/data-driven-deep-reinforcement-learning/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sat, 07 Dec 2019 18:03:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/data-driven-deep-reinforcement-learning/</guid>

					<description><![CDATA[<p>One of the primary factors behind the success of machine learning approaches in open world settings, such as <a href="https://arxiv.org/abs/1512.03385" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">image recognition</a> and <a href="https://arxiv.org/abs/1810.04805" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">natural language processing</a>, has been the ability of high-capacity deep neural network function approximators to learn generalizable models from large amounts of data. Deep reinforcement learning methods, however, require active online data collection, where the model actively interacts with its environment. This makes such methods hard to scale to complex real-world problems, where active data collection means that large datasets of experience must be collected for every experiment &#8211; this can be expensive and, for systems such as autonomous vehicles or robots, potentially <a href="https://arxiv.org/abs/1801.08757" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">unsafe</a>. In a number of domains of practical interest, such as autonomous driving, robotics, and games, there exist plentiful amounts of previously collected interaction data which, consists of informative behaviours that are a rich source of prior information. Deep RL algorithms that can utilize such prior datasets will not only scale to real-world problems, but will also lead to solutions that generalize substantially better. A <strong>data-driven</strong> paradigm for reinforcement learning will enable us to pre-train and deploy agents capable of sample-efficient learning in the real-world.</p>

<p>
<img src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575540575850_off_policy_teaser.png" width=""></p>

<p>In this work, we ask the following question: <em>Can deep RL algorithms effectively leverage prior collected offline data and learn without interaction with the environment?</em> We refer to this problem statement as <em>fully off-policy RL</em>, previously also called <em>batch RL</em> in literature. A class of deep RL algorithms, known as off-policy RL algorithms can, in principle, learn from previously collected data. Recent off-policy RL algorithms such as <a href="https://arxiv.org/abs/1812.05905" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Soft Actor-Critic</a> (SAC), <a href="https://arxiv.org/abs/1806.10293" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">QT-Opt</a>, and <a href="https://arxiv.org/abs/1710.02298" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Rainbow</a>, have demonstrated sample-efficient performance in a number of challenging domains such as robotic manipulation and atari games. However, all of these methods still require online data collection, and their ability to learn from fully off-policy data is limited in practice. In this work, we show why existing deep RL algorithms can fail in the fully off-policy setting. We then propose effective solutions to mitigate these issues.</p>

<!--more-->

<h2>Why can&#8217;t off-policy deep RL algorithms learn from static datasets?</h2>

<p>Let&#8217;s first study how state-of-the-art deep RL algorithms perform in the fully
off-policy setting. We choose the Soft Actor-critic (SAC) algorithm and
investigate its performance in the fully off-policy setting. Figure 1 shows the
training curve for SAC trained solely on varying amounts of previously
collected <em>expert demonstrations</em> for the HalfCheetah-v2 gym benchmark task.
Although the data demonstrates successful task completion, none of these runs
succeed, with corresponding Q-values diverging in some cases. At first glance,
this resembles <em><strong>overfitting</strong>,</em> as the evaluation performance deteriorates
with more training (green curve), but increasing the size of the static dataset
does not rectify the problem (orange (1e4 samples) vs green (1e5 samples)),
suggesting the issue is more complex.</p>

<p>
<img src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575540563406_returns_cheetah.png" width="600"><br><i>
Figure 1: Average Return and logarithm of the Q-value for HalfCheetah-v2 with varying amounts of expert data.
</i>
</p>

<p>Most off-policy RL methods in use today, including SAC used in Figure 1, are
based on approximate dynamic programming (though many also utilize importance
sampled policy gradients, for example <a href="http://auai.org/uai2019/proceedings/papers/440.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Liu et al.
(2019)</a>). The core
component of approximate dynamic programming in deep RL is the value function
or Q-function. The <em>optimal</em> Q-function  obeys the optimal Bellman
equation, given below:</p>

<p>Reinforcement learning then corresponds to minimizing the squared difference
between the left-hand side and right-hand side of this equation, also referred
to as the mean squared Bellman error (MSBE):</p>

<p>MSBE is minimized on transition samples in a dataset  generated
by a behavior policy . Although minimizing MSBE corresponds to a
supervised regression problem, the targets for this regression are themselves
derived from the current Q-function estimate.</p>

<p>We can understand the source of the instability shown in Figure 1 by examining
the form of the Bellman backup described above.  The targets are calculated by
maximizing the learned Q-values with respect to the action at the next state
() for Q-learning methods, or by computing an expected value under the
policy at the next state,  for actor-critic methods that maintain an explicit policy
alongside a Q-function. However, the Q-function estimator is only reliable on
action-inputs from the behavior policy , which is the training
distribution. As a result, naively maximizing the value may evaluate the
 estimator on actions that lie far outside of the training
distribution, resulting in pathological values that incur large absolute error
from the optimal desired Q-values. We refer to such actions as
out-of-distribution (OOD) actions. We call the error in the Q-values caused due
to backing up target values corresponding to OOD actions during backups as
bootstrapping error. A schematic of this problem is shown in Figure 2 below.</p>

<p>
<img src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575540995462_problem.png" width="600"><br><i>
Figure 2: Incorrectly high Q-values for OOD actions may be used for backups, leading to accumulation of error.
</i>
</p>

<p>Using an incorrect Q-value for computing the backup leads to accumulation of
error &#8211; minimizing MSBE perfectly leads to propagation of all imperfections in
the target values into the current Q-function estimator. We refer to this
process as accumulation or <strong>propagation</strong> of bootstrapping error over
iterations of training. The agent is unable to correct errors as it is unable
to gather ground truth return information by actively exploring the
environment. Accumulation of bootstrapping error can make Q-values diverge to
infinity (for example, blue and green curves in Figure 1), or not converge to
the correct values (for example, red curve in Figure 1), leading to final
performance that is substantially worse than even the average performance in
the dataset.</p>

<h2>Towards deep RL algorithms that learn in the presence of OOD actions</h2>

<p><em>How can we develop RL algorithms that learn from static data without being affected by OOD actions?</em> We will first review some of the existing approaches in literature towards solving this problem and then describe our recent work, BEAR which tackles this problem. There are broadly two classes of methods towards solving this problem.</p>

<ul><li><strong>Behavioral cloning (BC) based methods:</strong> When the static dataset is generated by an expert, one can utilize behavioral cloning to mimic the expert policy, as is done in imitation learning. In a generalized setting, where the behavior policy can be suboptimal but reward information is accessible, one can choose to only mimic a subset of good action decisions from the entire dataset. This is the idea behind prior works such as <a href="http://is.tuebingen.mpg.de/fileadmin/user_upload/files/publications/ICML2007-Peters_4493%5B0%5D.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reward-weighted regression</a> (RWR), <a href="https://www.aaai.org/ocs/index.php/AAAI/AAAI10/paper/viewFile/1851/2264" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">relative entropy policy search</a> (REPS), <a href="https://papers.nips.cc/paper/7866-exponentially-weighted-imitation-learning-for-batched-historical-data" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">MARWIL</a>, some very recent works such as <a href="https://openreview.net/forum?id=rke7geHtwH" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ABM</a>, and a recent work from our lab, called <a href="https://arxiv.org/abs/1910.00177" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">advantage weighted regression</a> (AWR) where we show that advantage-weighted form of behavioral cloning, which assigns higher likelihoods to demonstration actions that receive higher advantages, can also be used in such situations. Such a method trains only on actions observed in the dataset, hence avoids OOD actions completely.</li>
  <li><strong>Dynamic programming (DP) methods:</strong> Dynamic programming methods are appealing in fully off-policy RL scenarios because of their ability to pool information across trajectories, unlike BC-based methods that are implicitly constrained to lie in the vicinity of the best performing trajectory in the static dataset. For example, Q-iteration on a dataset consisting of all transitions in an MDP should return the optimal policy at convergence, however previously described BC-based methods may fail to recover optimality if the individual trajectories are highly suboptimal. Within this class, some recent work includes <a href="https://arxiv.org/abs/1812.02900" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">batch constrained Q-learning</a> (BCQ) that constrains the trained policy distribution to lie close to the behavior policy that generated the dataset. This is an optimal strategy when the static dataset is generated by an expert policy. However, this might be suboptimal if the data comes from an arbitrarily suboptimal policy. Other recent work, <a href="https://arxiv.org/abs/1512.08562" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">G-Learning</a>, <a href="https://arxiv.org/abs/1907.00456" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">KL-Control</a>, <a href="https://arxiv.org/abs/1911.11361" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BRAC</a> implements closeness to the behavior policy by solving a KL-constrained RL problem. <a href="https://arxiv.org/abs/1712.06924" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SPIBB</a> selectively constrains the learned policy to match the behavior policy in probability density on less frequent actions.</li>
</ul><p>The key question we pose in our work is: <strong>Which policies can be reliably used
for backups without backing up OOD actions?</strong> Once this question is answered,
the job of an RL algorithm reduces to picking the best policy in this set. In
our work, we provide a theoretical characterization of this set of policies and
use insights from theory to propose a practical dynamic programming based deep
RL algorithm called <a href="https://arxiv.org/abs/1906.00949" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BEAR</a> that learns from
purely static data.</p>

<h2>Bootstrapping Error Accumulation Reduction (BEAR)</h2>

<p>
<img src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575540392160_all_three_cases.png" width=""><br><i>
Figure 3: Illustration of support constraint (BEAR) (right) and distribution-matching constraint (middle).
</i>
</p>

<p>The key idea behind BEAR is to constrain the learned policy to lie <em>within the
support</em> (Figure 3, right) of the behavior policy distribution. This is in
contrast to distribution matching (Figure 3, middle) &#8211; BEAR does not constrain
the learned policy to be close in distribution to the behavior policy, but only
requires that the learned policy places non-zero probability mass on actions
with non-negligible behavior policy density. We refer to this as <strong>support
constraint.</strong> As an example, in a setting with a uniform-at-random behavior
policy, a support constraint allows dynamic programming to learn an optimal,
deterministic policy. However, a distribution-matching constraint will require
that the learned policy be highly stochastic (and thus not optimal), for
instance in Figure 3, middle, the learned policy is constrained to be one of
the stochastic purple policies, however in Figure 3, right, the learned policy
can be a (near-)deterministic yellow policy. For the readers interested in
theory, the theoretical insight behind this choice is that a support constraint
enables us to control error propagation by upper bounding
<a href="https://www.aaai.org/Papers/AAAI/2005/AAAI05-159.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">concentrability</a> under
the learned policy, while providing the capacity to reduce divergence from the
optimal policy.</p>

<p>How do we enforce that the learned policy satisfies the support constraint? In
practice, we use the sampled <a href="http://jmlr.csail.mit.edu/papers/v13/gretton12a.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Maximum Mean
Discrepancy</a> (MMD)
distance between actions as a measure of support divergence. Letting ,  and $k$ be any RBF kernel, we have:</p>

<p>A simple code snippet for computing MMD is shown below:</p>

<div><div><pre><code><span>def</span> <span>gaussian_kernel</span><span>(</span><span>x</span><span>,</span> <span>y</span><span>,</span> <span>sigma</span><span>=</span><span>0.1</span><span>):</span>
  <span>return</span> <span>exp</span><span>(</span><span>-</span><span>(</span><span>x</span> <span>-</span> <span>y</span><span>)</span><span>.</span><span>pow</span><span>(</span><span>2</span><span>)</span><span>.</span><span>sum</span><span>()</span> <span>/</span> <span>(</span><span>2</span> <span>*</span> <span>sigma</span><span>.</span><span>pow</span><span>(</span><span>2</span><span>)))</span>

<span>def</span> <span>compute_mmd</span><span>(</span><span>x</span><span>,</span> <span>y</span><span>):</span>
  <span>k_x_x</span> <span>=</span> <span>gaussian_kernel</span><span>(</span><span>x</span><span>,</span> <span>x</span><span>)</span>
  <span>k_x_y</span> <span>=</span> <span>gaussian_kernel</span><span>(</span><span>x</span><span>,</span> <span>y</span><span>)</span>
  <span>k_y_y</span> <span>=</span> <span>gaussian_kernel</span><span>(</span><span>y</span><span>,</span> <span>y</span><span>)</span>
  <span>return</span> <span>sqrt</span><span>(</span><span>k_x_x</span><span>.</span><span>mean</span><span>()</span> <span>+</span> <span>k_y_y</span><span>.</span><span>mean</span><span>()</span> <span>-</span> <span>2</span><span>*</span><span>k_x_y</span><span>.</span><span>mean</span><span>())</span>
</code></pre></div></div>

<p> is amenable to stochastic gradient-based training and we show
that computing  using only few samples from both
distributions  and  provides sufficient signal to quantify
differences in support but not in probability density, hence making it a
preferred measure for implementing the support constraint. To sum up, the new
(constrained-)policy improvement step in an actor-critic setup is given by:</p>

<p><strong>Support constraint vs Distribution-matching constraint</strong>
Some works, for example, <a href="https://arxiv.org/abs/1907.00456" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">KL-Control</a>, <a href="https://arxiv.org/abs/1911.11361" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BRAC</a>, <a href="https://arxiv.org/abs/1512.08562" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">G-Learning</a>, argue that using a distribution-matching constraint might suffice in such fully off-policy RL problems. In this section, we take a slight detour towards analyzing this choice. In particular, we provide an instance of an MDP where distribution-matching constraint might lead to arbitrarily suboptimal behavior while support matching does not suffer from this issue.</p>

<p>Consider the 1D-lineworld MDP shown in Figure 4 below. Two actions (left and right) are available to the agent at each state. The agent is tasked with reaching to the goal state , starting from state  and the corresponding per-step reward values are shown in Figure 4(a). The agent is only allowed to learn from behavior data generated by a policy that performs actions with probabilities described in Figure 4(b), and in particular, this behavior policy executes the suboptimal action at states in-between S and G with a high likelihood of 0.9, however, both actions  and  are in-distribution at all these states.</p>

<p>
<img src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575585059655_gridworld_separate_init.png" width=""><br><i>
Figure 4: Example 1D lineworld and the corresponding behavior policy.
</i>
</p>

<p>In Figure 5(a), we show that the learned policy with a distribution-matching constraint can be arbitrarily suboptimal, infact, the probability of reaching goal G by rolling out this policy is very small, and tends to 0 as the environment is made larger. However, in Figure 5(b), we show that a support constraint can recover an optimal policy with probability 1.</p>

<p>
<img src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575585070806_policies_learned.png" width=""><br><i>
Figure 5: Policies learned via distribution-matching and support-matching in the 1D lineworld shown in Figure 4.
</i>
</p>

<p>Why does distribution-matching fail here? Let us analyze the case when we use a penalty for distribution-matching. If the penalty is enforced tightly, then the agent will be forced to mainly execute the wrong action () in states between  and , leading to suboptimal behavior. However, if the penalty is not enforced tightly, with the intention of achieving a better policy than the behavior policy, the agent will perform backups using the OOD-action  at states to the left of , and these backups will eventually affect the Q-value at state . This phenomenon will lead to an incorrect Q-function, and hence a wrong policy &#8211; possibly, one that goes to the left starting from  instead of moving towards , as OOD-action backups combined with overestimation bias in Q-learning might make action  at state  look more preferable. Figure 6 shows that some states need a strong penalty/constraint (to prevent OOD backups) and the others require a weak penalty/constraint (to achieve optimality) for distribution-matching to work, however, this cannot be achieved via conventional distribution-matching approaches.</p>

<p>
<img src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575585085723_gridworld_analysis.png" width="600"><br><i>
Figure 6: Analysis of the strength of distribution-matching constraint needed at different states.
</i>
</p>

&#60;!--
**Support constraint vs Distribution-matching constraint**
Some works, for example, [KL-Control](https://arxiv.org/abs/1907.00456), [BRAC](https://arxiv.org/abs/1911.11361), [G-Learning](https://arxiv.org/abs/1512.08562), argue that using a distribution-matching constraint might suffice in such fully off-policy RL problems. In this section, we take a slight detour towards analyzing this choice. In particular, we provide an instance of an MDP where distribution-matching constraint might lead to arbitrarily suboptimal behavior while support matching does not suffer from this issue.

Consider the 1D-lineworld MDP shown in Figure 4 below. Two actions (left and right) are available to the agent at each state. The agent is tasked with reaching to the goal state $$G$$, starting from state $$S$$ and the corresponding per-step reward values are shown in Figure 4, left. The agent is only allowed to learn from behavior data generated by a policy that performs actions with probabilities described in Figure 3a, and in particular, this behavior policy executes the suboptimal action at states in-between S and G with a high likelihood of 0.9, however, both actions $$\leftarrow$$ and $$\rightarrow$$ are in-distribution at all these states. In Figure 4c, we show that the learned policy with a distribution-matching constraint can be arbitrarily suboptimal, infact, the probability of reaching goal G by rolling out this policy is very small, and tends to 0 as the environment is made larger. However, in Figure 4d, we show that a support constraint can recover an optimal policy with probability 1.

<p style="text-align:center">
<img src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575498994140_gridworld_example.png" width="">
<br />
<i>
Figure 4: Example 1D lineworld to illustrate the difference between support-constraint and distribution-matching constraint.
</i>
</p>

Why does this distribution-matching fail here? Let us imagine using a penalty
for distribution-matching. If the penalty is enforced tightly, then the agent
will be forced to mainly execute the wrong action in states between $$S$$ and
$$G$$, leading to suboptimal behavior. However, if the penalty is not enforced
tightly, with the objective of achieving more optimal behavior, this will lead
to backups from the OOD-action '$$\rightarrow$$' at states to the left of
$$S$$, which can lead to a completely wrong Q-function, and hence a wrong
policy -- for example, it can give rise to a policy that goes to the left
instead of the right. In Figure 4b, we demonstrate this issue by partitioning
states into groups that need a strong penalty/constraint and a group that needs
a weak penalty/constraint for distribution-matching to work, however, this
might not be achievable via conventional distribution-matching approaches.

--&#62;

<h2>So, how does BEAR perform in practice?</h2>

<p>In our experiments, we evaluated BEAR on three kinds of datasets generated by
&#8211; (1) a partially-trained medium-return policy, (2) a random low-return policy
and (3) an expert, high-return policy. (1) resembles the settings in practice
such as autonomous driving or robotics, where offline data is collected via
scripted policies for robotic grasping or consists of human driving data (which
may not be perfect) respectively. Such data is useful as it demonstrates
non-random, but still not optimal behavior and we expect training on offline
data to be most useful in this setting. Good performance on <em>both</em> (2) and (3)
demonstrates that the versatility of an algorithm to arbitrary dataset
compositions.</p>

<p>For each dataset composition, we compare BEAR to a number of baselines
including BC, BCQ, and deep Q-Learning from demonstrations
(<a href="https://arxiv.org/abs/1704.03732" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DQfD</a>). In general, we find that BEAR
outperforms the best performing baseline in setting (1), and BEAR is the only
algorithm capable successfully learning a better-than-dataset policy in both
(2) and (3). We show some learning curves below. BC or BCQ type methods usually
do not perform great with random data, partly because of the usage of a
distribution-matching constraint as described earlier.</p>

<p>
<img src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575194663346_Screen+Shot+2019-12-01+at+2.03.50+AM.png" width=""><br><i>
Figure 7: Performance on (1) medium-quality dataset: BEAR outperforms the best performing baseline.
</i>
</p>

<p>
<img src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575194824925_Screen+Shot+2019-12-01+at+2.06.48+AM.png" width="600"><br><i>
Figure 8: (3) Expert data: BEAR recovers the performance in the expert dataset, and performs similarly to other methods such as BC.
</i>
</p>

<p>
<img src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575194836848_Screen+Shot+2019-12-01+at+2.05.38+AM.png" width="600"><br><i>
Figure 9: (2) Random data: BEAR recovers better than dataset performance, and is close to the best performing algorithm (Naive RL).
</i>
</p>

<h2>Future Directions and Open Problems</h2>

<p>Most of the prior datasets for real-world problems such as <a href="https://www.robonet.wiki/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RoboNet</a> and <a href="https://github.com/TorchCraft/StarData" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Starcraft replays</a> consist of multimodal behavior generated by different users and robots. Hence, one of the next steps to look at is learning from a diverse mixtures of policies. How can we effectively learn policies from a static dataset that consists of a diverse range of behavior &#8211; possibly interaction from a diverse range of tasks, in the spirit of what we encounter in the real world? This question is mostly unanswered at the moment. Some very recent work, such as <a href="https://arxiv.org/abs/1907.04543" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">REM</a>, shows that simple modifications to existing distributional off-policy RL algorithms in the Atari domain can enable fully off-policy learning on entire interaction data generated from the training run of a separate DQN agent. However, the best solution for learning from a dataset generated by any arbitrary mixture of policies &#8211; which is more likely the case in practical problems &#8211; is unclear.</p>

<p>A rigorous theoretical characterization of the best achievable policy as a function of a given dataset is also an open problem. In our paper, we analyze this question by looking at typically used assumptions of bounded concentrability in the error and convergence analysis of Fitted Q-iteration. Which other assumptions can be applied to analyze this problem? And which of these assumptions is least restrictive and practically feasible? What is the theoretical optimum of what can be achieved solely by offline training?</p>

<p>We hope that our work, BEAR, takes us a step closer to effectively leveraging the most out of prior datasets in an RL algorithm. A <strong>data-driven</strong> paradigm of RL where one could (pre-)train RL algorithms with large amounts of prior data will enable us to go beyond the active exploration bottleneck, thus giving us agents that can be deployed and keep learning continuously in the real world.</p>

<hr><p>This blog post is based on the our recent paper:</p>

<ul><li><strong>Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction</strong><br>
  Aviral Kumar*, Justin Fu*, George Tucker, Sergey Levine<br><em>In Advances in Neural Information Processing Systems, 2019</em></li>
</ul><p>The <a href="https://arxiv.org/abs/1906.00949" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a> and code are available
<a href="https://github.com/aviralkumar2907/BEAR" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">online</a> and a slide-deck explaining
the algorithm is available
<a href="https://sites.google.com/view/bear-off-policyrl" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>. <em>I would like to thank
Sergey Levine for his valuable feedback on earlier versions of this blog post.</em></p>]]></description>
										<content:encoded><![CDATA[<p><strong>By Aviral Kumar</strong></p>
<p>One of the primary factors behind the success of machine learning approaches in open world settings, such as <a href="https://arxiv.org/abs/1512.03385" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">image recognition</a> and <a href="https://arxiv.org/abs/1810.04805" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">natural language processing</a>, has been the ability of high-capacity deep neural network function approximators to learn generalizable models from large amounts of data. Deep reinforcement learning methods, however, require active online data collection, where the model actively interacts with its environment. This makes such methods hard to scale to complex real-world problems, where active data collection means that large datasets of experience must be collected for every experiment – this can be expensive and, for systems such as autonomous vehicles or robots, potentially <a href="https://arxiv.org/abs/1801.08757" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">unsafe</a>. In a number of domains of practical interest, such as autonomous driving, robotics, and games, there exist plentiful amounts of previously collected interaction data which, consists of informative behaviours that are a rich source of prior information. Deep RL algorithms that can utilize such prior datasets will not only scale to real-world problems, but will also lead to solutions that generalize substantially better. A <strong>data-driven</strong> paradigm for reinforcement learning will enable us to pre-train and deploy agents capable of sample-efficient learning in the real-world.</p>
<p> <span id="more-156532"></span></p>
<img decoding="async" src="https://robohub.org/wp-content/uploads/2019/12/BAIRRL.png" alt="" width="900" height="302" class="aligncenter size-full wp-image-156676" srcset="https://robohub.org/wp-content/uploads/2019/12/BAIRRL.png 900w, https://robohub.org/wp-content/uploads/2019/12/BAIRRL-425x143.png 425w, https://robohub.org/wp-content/uploads/2019/12/BAIRRL-768x258.png 768w" sizes="(max-width: 900px) 100vw, 900px" />
<p>In this work, we ask the following question: <em>Can deep RL algorithms effectively leverage prior collected offline data and learn without interaction with the environment?</em> We refer to this problem statement as <em>fully off-policy RL</em>, previously also called <em>batch RL</em> in literature. A class of deep RL algorithms, known as off-policy RL algorithms can, in principle, learn from previously collected data. Recent off-policy RL algorithms such as <a href="https://arxiv.org/abs/1812.05905" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Soft Actor-Critic</a> (SAC), <a href="https://arxiv.org/abs/1806.10293" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">QT-Opt</a>, and <a href="https://arxiv.org/abs/1710.02298" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Rainbow</a>, have demonstrated sample-efficient performance in a number of challenging domains such as robotic manipulation and atari games. However, all of these methods still require online data collection, and their ability to learn from fully off-policy data is limited in practice. In this work, we show why existing deep RL algorithms can fail in the fully off-policy setting. We then propose effective solutions to mitigate these issues.</p>
<p>  <!--more-->  </p>
<h2 id="why-cant-off-policy-deep-rl-algorithms-learn-from-static-datasets">Why can’t off-policy deep RL algorithms learn from static datasets?</h2>
<p>Let’s first study how state-of-the-art deep RL algorithms perform in the fully off-policy setting. We choose the Soft Actor-critic (SAC) algorithm and investigate its performance in the fully off-policy setting. Figure 1 shows the training curve for SAC trained solely on varying amounts of previously collected <em>expert demonstrations</em> for the HalfCheetah-v2 gym benchmark task. Although the data demonstrates successful task completion, none of these runs succeed, with corresponding Q-values diverging in some cases. At first glance, this resembles <em><strong>overfitting</strong>,</em> as the evaluation performance deteriorates with more training (green curve), but increasing the size of the static dataset does not rectify the problem (orange (1e4 samples) vs green (1e5 samples)), suggesting the issue is more complex.</p>
<p style="text-align:center;"> <img decoding="async" src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575540563406_returns_cheetah.png" width="600" /> <br /> <i> Figure 1: Average Return and logarithm of the Q-value for HalfCheetah-v2 with varying amounts of expert data. </i> </p>
<p>Most off-policy RL methods in use today, including SAC used in Figure 1, are based on approximate dynamic programming (though many also utilize importance sampled policy gradients, for example <a href="http://auai.org/uai2019/proceedings/papers/440.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Liu et al. (2019)</a>). The core component of approximate dynamic programming in deep RL is the value function or Q-function. The <em>optimal</em> Q-function <script type="math/tex">Q^*</script> obeys the optimal Bellman equation, given below:</p>
<p>  <script type="math/tex; mode=display">Q^* = \mathcal{T}^* Q^*  \;\;\;  \mbox{;} \;\;\;  (\mathcal{T}^* \hat{Q})(s, a) := R(s, a) + \gamma \mathbb{E}_{T(s'|s,a)}[\max_{a'}\hat{Q}(s', a')]</script>  </p>
<p>Reinforcement learning then corresponds to minimizing the squared difference between the left-hand side and right-hand side of this equation, also referred to as the mean squared Bellman error (MSBE):</p>
<p>  <script type="math/tex; mode=display">Q := \arg \min_{\hat{Q}} \mathbb{E}_{s \sim \mathcal{D}, a \sim \beta(a|s)} \left[ (\hat{Q}(s, a) - (\mathcal{T}^* \hat{Q})(s, a))^2 \right]</script>  </p>
<p>MSBE is minimized on transition samples in a dataset <script type="math/tex">\mathcal{D}</script> generated by a behavior policy <script type="math/tex">\beta(a|s)</script>. Although minimizing MSBE corresponds to a supervised regression problem, the targets for this regression are themselves derived from the current Q-function estimate.</p>
<p>We can understand the source of the instability shown in Figure 1 by examining the form of the Bellman backup described above.  The targets are calculated by maximizing the learned Q-values with respect to the action at the next state (<script type="math/tex">s'</script>) for Q-learning methods, or by computing an expected value under the policy at the next state, <script type="math/tex">\mathbb{E}_{s' \sim T(s'|s, a), a' \sim \pi(a'|s')} [\hat{Q}(s', a')]</script> for actor-critic methods that maintain an explicit policy alongside a Q-function. However, the Q-function estimator is only reliable on action-inputs from the behavior policy <script type="math/tex">\beta</script>, which is the training distribution. As a result, naively maximizing the value may evaluate the <script type="math/tex">\hat{Q}</script> estimator on actions that lie far outside of the training distribution, resulting in pathological values that incur large absolute error from the optimal desired Q-values. We refer to such actions as out-of-distribution (OOD) actions. We call the error in the Q-values caused due to backing up target values corresponding to OOD actions during backups as bootstrapping error. A schematic of this problem is shown in Figure 2 below.</p>
<p style="text-align:center;"> <img decoding="async" src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575540995462_problem.png" width="600" /> <br /> <i> Figure 2: Incorrectly high Q-values for OOD actions may be used for backups, leading to accumulation of error. </i> </p>
<p>Using an incorrect Q-value for computing the backup leads to accumulation of error – minimizing MSBE perfectly leads to propagation of all imperfections in the target values into the current Q-function estimator. We refer to this process as accumulation or <strong>propagation</strong> of bootstrapping error over iterations of training. The agent is unable to correct errors as it is unable to gather ground truth return information by actively exploring the environment. Accumulation of bootstrapping error can make Q-values diverge to infinity (for example, blue and green curves in Figure 1), or not converge to the correct values (for example, red curve in Figure 1), leading to final performance that is substantially worse than even the average performance in the dataset.</p>
<h2 id="towards-deep-rl-algorithms-that-learn-in-the-presence-of-ood-actions">Towards deep RL algorithms that learn in the presence of OOD actions</h2>
<p><em>How can we develop RL algorithms that learn from static data without being affected by OOD actions?</em> We will first review some of the existing approaches in literature towards solving this problem and then describe our recent work, BEAR which tackles this problem. There are broadly two classes of methods towards solving this problem.</p>
<p><strong>Behavioral cloning (BC) based methods:</strong> When the static dataset is generated by an expert, one can utilize behavioral cloning to mimic the expert policy, as is done in imitation learning. In a generalized setting, where the behavior policy can be suboptimal but reward information is accessible, one can choose to only mimic a subset of good action decisions from the entire dataset. This is the idea behind prior works such as <a href="http://is.tuebingen.mpg.de/fileadmin/user_upload/files/publications/ICML2007-Peters_4493%5B0%5D.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reward-weighted regression</a> (RWR), <a href="https://www.aaai.org/ocs/index.php/AAAI/AAAI10/paper/viewFile/1851/2264" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">relative entropy policy search</a> (REPS), <a href="https://papers.nips.cc/paper/7866-exponentially-weighted-imitation-learning-for-batched-historical-data" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">MARWIL</a>, some very recent works such as <a href="https://openreview.net/forum?id=rke7geHtwH" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ABM</a>, and a recent work from our lab, called <a href="https://arxiv.org/abs/1910.00177" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">advantage weighted regression</a> (AWR) where we show that advantage-weighted form of behavioral cloning, which assigns higher likelihoods to demonstration actions that receive higher advantages, can also be used in such situations. Such a method trains only on actions observed in the dataset, hence avoids OOD actions completely.</p>
<p><strong>Dynamic programming (DP) methods:</strong> Dynamic programming methods are appealing in fully off-policy RL scenarios because of their ability to pool information across trajectories, unlike BC-based methods that are implicitly constrained to lie in the vicinity of the best performing trajectory in the static dataset. For example, Q-iteration on a dataset consisting of all transitions in an MDP should return the optimal policy at convergence, however previously described BC-based methods may fail to recover optimality if the individual trajectories are highly suboptimal. Within this class, some recent work includes <a href="https://arxiv.org/abs/1812.02900" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">batch constrained Q-learning</a> (BCQ) that constrains the trained policy distribution to lie close to the behavior policy that generated the dataset. This is an optimal strategy when the static dataset is generated by an expert policy. However, this might be suboptimal if the data comes from an arbitrarily suboptimal policy. Other recent work, <a href="https://arxiv.org/abs/1512.08562" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">G-Learning</a>, <a href="https://arxiv.org/abs/1907.00456" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">KL-Control</a>, <a href="https://arxiv.org/abs/1911.11361" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BRAC</a> implements closeness to the behavior policy by solving a KL-constrained RL problem. <a href="https://arxiv.org/abs/1712.06924" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SPIBB</a> selectively constrains the learned policy to match the behavior policy in probability density on less frequent actions.</p>
<p>The key question we pose in our work is: <strong>Which policies can be reliably used for backups without backing up OOD actions?</strong> Once this question is answered, the job of an RL algorithm reduces to picking the best policy in this set. In our work, we provide a theoretical characterization of this set of policies and use insights from theory to propose a practical dynamic programming based deep RL algorithm called <a href="https://arxiv.org/abs/1906.00949" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BEAR</a> that learns from purely static data.</p>
<h2 id="bootstrapping-error-accumulation-reduction-bear">Bootstrapping Error Accumulation Reduction (BEAR)</h2>
<p style="text-align:center;"> <img decoding="async" src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575540392160_all_three_cases.png" width="" /> <br /> <i> Figure 3: Illustration of support constraint (BEAR) (right) and distribution-matching constraint (middle). </i> </p>
<p>The key idea behind BEAR is to constrain the learned policy to lie <em>within the support</em> (Figure 3, right) of the behavior policy distribution. This is in contrast to distribution matching (Figure 3, middle) – BEAR does not constrain the learned policy to be close in distribution to the behavior policy, but only requires that the learned policy places non-zero probability mass on actions with non-negligible behavior policy density. We refer to this as <strong>support constraint.</strong> As an example, in a setting with a uniform-at-random behavior policy, a support constraint allows dynamic programming to learn an optimal, deterministic policy. However, a distribution-matching constraint will require that the learned policy be highly stochastic (and thus not optimal), for instance in Figure 3, middle, the learned policy is constrained to be one of the stochastic purple policies, however in Figure 3, right, the learned policy can be a (near-)deterministic yellow policy. For the readers interested in theory, the theoretical insight behind this choice is that a support constraint enables us to control error propagation by upper bounding <a href="https://www.aaai.org/Papers/AAAI/2005/AAAI05-159.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">concentrability</a> under the learned policy, while providing the capacity to reduce divergence from the optimal policy.</p>
<p>How do we enforce that the learned policy satisfies the support constraint? In practice, we use the sampled <a href="http://jmlr.csail.mit.edu/papers/v13/gretton12a.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Maximum Mean Discrepancy</a> (MMD) distance between actions as a measure of support divergence. Letting <script type="math/tex">X = \{x_1, \cdots, x_n\}</script>, <script type="math/tex">Y = \{y_1, \cdots, y_n\}</script> and $k$ be any RBF kernel, we have:</p>
<p>  <script type="math/tex; mode=display">\text{MMD}^2(X, Y) = \frac{1}{n^2} \sum_{i, i'} k(x_i, x_{i'}) - \frac{2}{nm} \sum_{i, j} k(x_i, y_j) + \frac{1}{m^2} \sum_{j, j'} k(y_j, y_{j'}).</script>  </p>
<p>A simple code snippet for computing MMD is shown below:</p>
<img decoding="async" src="https://robohub.org/wp-content/uploads/2019/12/BAIRcode.png" alt="" width="1644" height="444" class="aligncenter size-full wp-image-156670" srcset="https://robohub.org/wp-content/uploads/2019/12/BAIRcode.png 1644w, https://robohub.org/wp-content/uploads/2019/12/BAIRcode-425x115.png 425w, https://robohub.org/wp-content/uploads/2019/12/BAIRcode-768x207.png 768w, https://robohub.org/wp-content/uploads/2019/12/BAIRcode-1024x277.png 1024w" sizes="(max-width: 1644px) 100vw, 1644px" />
<p><script type="math/tex">\text{MMD}</script> is amenable to stochastic gradient-based training and we show that computing <script type="math/tex">\text{MMD}(P, Q)</script> using only few samples from both distributions <script type="math/tex">P</script> and <script type="math/tex">Q</script> provides sufficient signal to quantify differences in support but not in probability density, hence making it a preferred measure for implementing the support constraint. To sum up, the new (constrained-)policy improvement step in an actor-critic setup is given by:</p>
<p>  <script type="math/tex; mode=display">\pi_\phi := \max_{\pi \in \Delta_{|S|}} \mathbb{E}_{s \sim \mathcal{D}} \mathbb{E}_{a \sim \pi(\cdot|s)} \left[ Q_\theta(s, a)\right] \quad \mbox{s.t.}  \quad \mathbb{E}_{s \sim \mathcal{D}} [\text{MMD}(\beta(\cdot|s), \pi(\cdot|s))] \leq \varepsilon</script>  </p>
<p><strong>Support constraint vs Distribution-matching constraint</strong> Some works, for example, <a href="https://arxiv.org/abs/1907.00456" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">KL-Control</a>, <a href="https://arxiv.org/abs/1911.11361" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BRAC</a>, <a href="https://arxiv.org/abs/1512.08562" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">G-Learning</a>, argue that using a distribution-matching constraint might suffice in such fully off-policy RL problems. In this section, we take a slight detour towards analyzing this choice. In particular, we provide an instance of an MDP where distribution-matching constraint might lead to arbitrarily suboptimal behavior while support matching does not suffer from this issue.</p>
<p>Consider the 1D-lineworld MDP shown in Figure 4 below. Two actions (left and right) are available to the agent at each state. The agent is tasked with reaching to the goal state <script type="math/tex">G</script>, starting from state <script type="math/tex">S</script> and the corresponding per-step reward values are shown in Figure 4(a). The agent is only allowed to learn from behavior data generated by a policy that performs actions with probabilities described in Figure 4(b), and in particular, this behavior policy executes the suboptimal action at states in-between S and G with a high likelihood of 0.9, however, both actions <script type="math/tex">\leftarrow</script> and <script type="math/tex">\rightarrow</script> are in-distribution at all these states.</p>
<p style="text-align:center;"> <img decoding="async" src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575585059655_gridworld_separate_init.png" width="" /> <br /> <i> Figure 4: Example 1D lineworld and the corresponding behavior policy. </i> </p>
<p>In Figure 5(a), we show that the learned policy with a distribution-matching constraint can be arbitrarily suboptimal, infact, the probability of reaching goal G by rolling out this policy is very small, and tends to 0 as the environment is made larger. However, in Figure 5(b), we show that a support constraint can recover an optimal policy with probability 1.</p>
<p style="text-align:center;"> <img decoding="async" src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575585070806_policies_learned.png" width="" /> <br /> <i> Figure 5: Policies learned via distribution-matching and support-matching in the 1D lineworld shown in Figure 4. </i> </p>
<p>Why does distribution-matching fail here? Let us analyze the case when we use a penalty for distribution-matching. If the penalty is enforced tightly, then the agent will be forced to mainly execute the wrong action (<script type="math/tex">\leftarrow</script>) in states between <script type="math/tex">S</script> and <script type="math/tex">G</script>, leading to suboptimal behavior. However, if the penalty is not enforced tightly, with the intention of achieving a better policy than the behavior policy, the agent will perform backups using the OOD-action <script type="math/tex">\rightarrow</script> at states to the left of <script type="math/tex">S</script>, and these backups will eventually affect the Q-value at state <script type="math/tex">S</script>. This phenomenon will lead to an incorrect Q-function, and hence a wrong policy – possibly, one that goes to the left starting from <script type="math/tex">S</script> instead of moving towards <script type="math/tex">G</script>, as OOD-action backups combined with overestimation bias in Q-learning might make action <script type="math/tex">\leftarrow</script> at state <script type="math/tex">S</script> look more preferable. Figure 6 shows that some states need a strong penalty/constraint (to prevent OOD backups) and the others require a weak penalty/constraint (to achieve optimality) for distribution-matching to work, however, this cannot be achieved via conventional distribution-matching approaches.</p>
<p style="text-align:center;"> <img decoding="async" src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575585085723_gridworld_analysis.png" width="600" /> <br /> <i> Figure 6: Analysis of the strength of distribution-matching constraint needed at different states. </i> </p>
<p>  <!-- **Support constraint vs Distribution-matching constraint** Some works, for example, [KL-Control](https://arxiv.org/abs/1907.00456), [BRAC](https://arxiv.org/abs/1911.11361), [G-Learning](https://arxiv.org/abs/1512.08562), argue that using a distribution-matching constraint might suffice in such fully off-policy RL problems. In this section, we take a slight detour towards analyzing this choice. In particular, we provide an instance of an MDP where distribution-matching constraint might lead to arbitrarily suboptimal behavior while support matching does not suffer from this issue.  Consider the 1D-lineworld MDP shown in Figure 4 below. Two actions (left and right) are available to the agent at each state. The agent is tasked with reaching to the goal state $$G$$, starting from state $$S$$ and the corresponding per-step reward values are shown in Figure 4, left. The agent is only allowed to learn from behavior data generated by a policy that performs actions with probabilities described in Figure 3a, and in particular, this behavior policy executes the suboptimal action at states in-between S and G with a high likelihood of 0.9, however, both actions $$\leftarrow$$ and $$\rightarrow$$ are in-distribution at all these states. In Figure 4c, we show that the learned policy with a distribution-matching constraint can be arbitrarily suboptimal, infact, the probability of reaching goal G by rolling out this policy is very small, and tends to 0 as the environment is made larger. However, in Figure 4d, we show that a support constraint can recover an optimal policy with probability 1.  

<p style="text-align:center;"> <img decoding="async" src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575498994140_gridworld_example.png" width=""> <br /> <i> Figure 4: Example 1D lineworld to illustrate the difference between support-constraint and distribution-matching constraint. </i> </p>

  Why does this distribution-matching fail here? Let us imagine using a penalty for distribution-matching. If the penalty is enforced tightly, then the agent will be forced to mainly execute the wrong action in states between $$S$$ and $$G$$, leading to suboptimal behavior. However, if the penalty is not enforced tightly, with the objective of achieving more optimal behavior, this will lead to backups from the OOD-action '$$\rightarrow$$' at states to the left of $$S$$, which can lead to a completely wrong Q-function, and hence a wrong policy -- for example, it can give rise to a policy that goes to the left instead of the right. In Figure 4b, we demonstrate this issue by partitioning states into groups that need a strong penalty/constraint and a group that needs a weak penalty/constraint for distribution-matching to work, however, this might not be achievable via conventional distribution-matching approaches.  -->  </p>
<h2 id="so-how-does-bear-perform-in-practice">So, how does BEAR perform in practice?</h2>
<p>In our experiments, we evaluated BEAR on three kinds of datasets generated by – (1) a partially-trained medium-return policy, (2) a random low-return policy and (3) an expert, high-return policy. (1) resembles the settings in practice such as autonomous driving or robotics, where offline data is collected via scripted policies for robotic grasping or consists of human driving data (which may not be perfect) respectively. Such data is useful as it demonstrates non-random, but still not optimal behavior and we expect training on offline data to be most useful in this setting. Good performance on <em>both</em> (2) and (3) demonstrates that the versatility of an algorithm to arbitrary dataset compositions.</p>
<p>For each dataset composition, we compare BEAR to a number of baselines including BC, BCQ, and deep Q-Learning from demonstrations (<a href="https://arxiv.org/abs/1704.03732" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DQfD</a>). In general, we find that BEAR outperforms the best performing baseline in setting (1), and BEAR is the only algorithm capable successfully learning a better-than-dataset policy in both (2) and (3). We show some learning curves below. BC or BCQ type methods usually do not perform great with random data, partly because of the usage of a distribution-matching constraint as described earlier.</p>
<p style="text-align:center;"> <img decoding="async" src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575194663346_Screen+Shot+2019-12-01+at+2.03.50+AM.png" width="" /> <br /> <i> Figure 7: Performance on (1) medium-quality dataset: BEAR outperforms the best performing baseline. </i> </p>
<p style="text-align:center;"> <img decoding="async" src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575194824925_Screen+Shot+2019-12-01+at+2.06.48+AM.png" width="600" /> <br /> <i> Figure 8: (3) Expert data: BEAR recovers the performance in the expert dataset, and performs similarly to other methods such as BC. </i> </p>
<p style="text-align:center;"> <img decoding="async" src="https://paper-attachments.dropbox.com/s_03D8A88577B961181603AE5EDBD4A511CD8E828E7651B8AA640A61950DAB9783_1575194836848_Screen+Shot+2019-12-01+at+2.05.38+AM.png" width="600" /> <br /> <i> Figure 9: (2) Random data: BEAR recovers better than dataset performance, and is close to the best performing algorithm (Naive RL). </i> </p>
<h2 id="future-directions-and-open-problems">Future Directions and Open Problems</h2>
<p>Most of the prior datasets for real-world problems such as <a href="https://www.robonet.wiki/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RoboNet</a> and <a href="https://github.com/TorchCraft/StarData" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Starcraft replays</a> consist of multimodal behavior generated by different users and robots. Hence, one of the next steps to look at is learning from a diverse mixtures of policies. How can we effectively learn policies from a static dataset that consists of a diverse range of behavior – possibly interaction from a diverse range of tasks, in the spirit of what we encounter in the real world? This question is mostly unanswered at the moment. Some very recent work, such as <a href="https://arxiv.org/abs/1907.04543" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">REM</a>, shows that simple modifications to existing distributional off-policy RL algorithms in the Atari domain can enable fully off-policy learning on entire interaction data generated from the training run of a separate DQN agent. However, the best solution for learning from a dataset generated by any arbitrary mixture of policies – which is more likely the case in practical problems – is unclear.</p>
<p>A rigorous theoretical characterization of the best achievable policy as a function of a given dataset is also an open problem. In our paper, we analyze this question by looking at typically used assumptions of bounded concentrability in the error and convergence analysis of Fitted Q-iteration. Which other assumptions can be applied to analyze this problem? And which of these assumptions is least restrictive and practically feasible? What is the theoretical optimum of what can be achieved solely by offline training?</p>
<p>We hope that our work, BEAR, takes us a step closer to effectively leveraging the most out of prior datasets in an RL algorithm. A <strong>data-driven</strong> paradigm of RL where one could (pre-)train RL algorithms with large amounts of prior data will enable us to go beyond the active exploration bottleneck, thus giving us agents that can be deployed and keep learning continuously in the real world.</p>
<hr />
<p>This blog post is based on the our recent paper:</p>
<ul>
<li><strong>Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction</strong><br />   Aviral Kumar*, Justin Fu*, George Tucker, Sergey Levine<br />   <em>In Advances in Neural Information Processing Systems, 2019</em></li>
</ul>
<p>The <a href="https://arxiv.org/abs/1906.00949" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a> and code are available <a href="https://github.com/aviralkumar2907/BEAR" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">online</a> and a slide-deck explaining the algorithm is available <a href="https://sites.google.com/view/bear-off-policyrl" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>. <em>I would like to thank Sergey Levine for his valuable feedback on earlier versions of this blog post. This article was initially published on the BAIR blog, and appears here with the authors’ permission.</em></p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>RoboNet: A dataset for large-scale multi-robot learning</title>
		<link>https://robohub.org/robonet-a-dataset-for-large-scale-multi-robot-learning/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sat, 07 Dec 2019 17:11:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/robonet-a-dataset-for-large-scale-multi-robot-learning/</guid>

					<description><![CDATA[This post is cross-listed at the
SAIL Blog and the
CMU ML blog.

In the last decade, we’ve seen learning-based systems provide transformative
solutions for a wide range of perception and reasoning problems, from
recognizing objects in
images
to recogni...]]></description>
										<content:encoded><![CDATA[<p><strong>By Sudeep Dasari</strong></p>
<p><i>This post is cross-listed at the <a href="http://ai.stanford.edu/blog/robonet" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SAIL Blog</a> and the <a href="https://blog.ml.cmu.edu/2019/11/26/robonet/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">CMU ML blog</a></i>.</p>
<p>In the last decade, we’ve seen learning-based systems provide transformative solutions for a wide range of perception and reasoning problems, from <a href="https://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recognizing objects in images</a> to <a href="https://ai.googleblog.com/2019/10/exploring-massively-multilingual.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recognizing and translating human speech</a>. Recent progress in deep reinforcement learning (i.e. integrating deep neural networks into reinforcement learning systems) suggests that the same kind of success could be realized in automated decision making domains. If fruitful, this line of work could allow learning-based systems to tackle active control tasks, such as robotics and autonomous driving, alongside the passive perception tasks to which they have already been successfully applied.</p>
<p>  <span id="more-155738"></span>  </p>
<p>While deep reinforcement learning methods &#8211; like <a href="https://bair.berkeley.edu/blog/2018/12/14/sac/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Soft Actor Critic</a> &#8211; can learn impressive motor skills, they are challenging to train on large and broad data that is not from the target environment. In contrast, the success of deep networks in fields like computer vision was arguably predicated just as much on large datasets, such as <a href="http://www.image-net.org/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ImageNet</a>, as it was on large neural network architectures. This suggests that applying data-driven methods to robotics will require not just the development of strong reinforcement learning methods, but also access to large and diverse datasets for robotics. Not only can large datasets enable models that generalize effectively, but they can also be used to <em>pre-train</em> models that can then be adapted to more specialized tasks using much more modest datasets. Indeed, “ImageNet pre-training” has become a default approach for tackling diverse tasks with small or medium datasets &#8211; like <a href="https://medium.com/geoai/reconstructing-3d-buildings-from-aerial-lidar-with-ai-details-6a81cb3079c0" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">3D building reconstruction</a>. Can the same kind of approach be adopted to enable broad generalization and transfer in active control domains, such as robotics?</p>
<p>Unfortunately, the design and adoption of large datasets in reinforcement learning and robotics has proven challenging. Since every robotics lab has their own hardware and experimental set-up, it is not apparent how to move towards an “ImageNet-scale” dataset for robotics that is useful for the entire research community. Hence, we propose to collect data across multiple different settings, including from varying camera viewpoints, varying environments, and even varying robot platforms. Motivated by the success of large-scale data-driven learning, we created RoboNet, an extensible and diverse dataset of robot interaction collected across <a href="https://bair.berkeley.edu/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">four</a> <a href="https://ai.stanford.edu/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">different</a> <a href="https://www.grasp.upenn.edu/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">research</a> <a href="https://ai.google/research/teams/brain/robotics/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">labs</a>. The collaborative nature of this work allows us to easily capture diverse data in various lab settings across a wide variety of objects, robotic hardware, and camera viewpoints. Finally, we find that pre-training on RoboNet offers substantial performance gains compared to training from scratch in entirely new environments.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/hypothesis.png" width="" /> <br /> <i> Our goal is to pre-train reinforcement learning models on a sufficiently diverse dataset and then transfer knowledge (either zero-shot or with fine-tuning) to a different test environment. </i> </p>
<h1 id="collecting-robonet">Collecting RoboNet</h1>
<p>RoboNet consists of 15 million video frames, collected by different robots interacting with different objects in a table-top setting. Every frame includes the image recorded by the robot’s camera, arm pose, force sensor readings, and gripper state. The collection environment, including the camera view, the appearance of the table or bin, and the objects in front of the robot are varied between trials. Since collection is entirely autonomous, large amounts can be cheaply collected across multiple institutions. A sample of RoboNet along with data statistics is shown below:</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/tile.gif" height="260" width="" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/data.png" height="260" width="" /> <br /> <i> A sample of data from RoboNet alongside a summary of the current dataset. Note that any GIF compression artifacts in this animation are not present in the dataset itself. </i> </p>
<h1 id="how-can-we-use-robonet">How can we use RoboNet?</h1>
<p>After collecting a diverse dataset, we experimentally investigate how it can be used to enable <em>general</em> skill learning that transfers to new environments. First, we pre-train <a href="https://alexlee-gk.github.io/video_prediction/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">visual dynamics models</a> on a subset of data from RoboNet, and then fine-tune them to work in an unseen test environment using a small amount of new data. The constructed test environments (one of which is visualized below) all include different lab settings, new cameras and viewpoints, held-out robots, and novel objects purchased after data collection concluded.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/test.jpg" height="400" width="" /> <br /> <br />
<i> Example test environment constructed in a new lab, with a temporary uncalibrated camera, and a new Baxter robot. Note that while Baxters are present in RoboNet that data is <i>not</i> included during model pre-training. </i> </p>
<p>After tuning, we deploy the learned dynamics models in the test environment to perform control tasks &#8211; like picking and placing objects &#8211; using the <a href="https://bair.berkeley.edu/blog/2018/11/30/visual-rl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">visual foresight</a> model based reinforcement learning algorithm. Below are example control tasks executed in various test environments.</p>
<p>  <!-- For now using the CMU links, if they expire just use our static ones below. --> </p>
<div class="containerWide">
<div class="photosWide">     <img decoding="async" class="imageWide" src="https://bair.berkeley.edu/static/blog/robonet/align_tshirt.gif" width="" height="185" />     <!-- <img decoding="async" class="imageWide" src="https://blog.ml.cmu.edu/wp-content/uploads/2019/11/align_tshirt.gif" width="" height="185"> -->     <span class="wordWide"><br />
<i>Kuka can align shirts next to the others</i></span>   </div>
<div class="photosWide">       <img decoding="async" class="imageWide" src="https://bair.berkeley.edu/static/blog/robonet/color_stripe_front.gif" width="" height="185" />       <!-- <img decoding="async" class="imageWide" src="https://blog.ml.cmu.edu/wp-content/uploads/2019/11/color_stripe_front.gif" width="" height="185"> -->       <span class="wordWide"><br />
<i>Baxter can sweep the table with cloth</i></span>   </div>
<div class="photosWide">       <img decoding="async" class="imageWide" src="https://bair.berkeley.edu/static/blog/robonet/marker_marker.gif" width="" height="185" />       <!-- <img decoding="async" class="imageWide" src="https://blog.ml.cmu.edu/wp-content/uploads/2019/11/marker_marker.gif" width="" height="185"> -->       <span class="wordWide"><br />
<i>Franka can grasp and reposition the markers</i></span>   </div>
</p></div>
<div class="containerWide">
<div class="photosWide">     <img decoding="async" class="imageWide" src="https://bair.berkeley.edu/static/blog/robonet/move_plate.gif" width="" height="185" />     <!-- <img decoding="async" class="imageWide" src="https://blog.ml.cmu.edu/wp-content/uploads/2019/11/move_plate.gif" width="" height="185"> -->     <span class="wordWide"><br />
<i> Kuka can move the plate to the edge of the table</i></span>   </div>
<div class="photosWide">       <img decoding="async" class="imageWide" src="https://bair.berkeley.edu/static/blog/robonet/socks_right.gif" width="" height="185" />       <!-- <img decoding="async" class="imageWide" src="https://blog.ml.cmu.edu/wp-content/uploads/2019/11/socks_right.gif" width="" height="185"> -->       <span class="wordWide"><br />
<i>Baxter can pick up and reposition socks </i></span>   </div>
<div class="photosWide">       <img decoding="async" class="imageWide" src="https://bair.berkeley.edu/static/blog/robonet/towel_stack.gif" width="" height="185" />       <!-- <img decoding="async" class="imageWide" src="https://blog.ml.cmu.edu/wp-content/uploads/2019/11/towel_stack.gif" width="" height="185"> -->       <span class="wordWide"><br />
<i>Franka can stack the towel on the pile</i></span>   </div>
</p></div>
<p style="text-align:center;"> <br />
<i> Here you can see examples of visual foresight fine-tuned to perform basic control tasks in three entirely different environments. For the experiments, the target robot and environment was subtracted from RoboNet during pre-training. Fine-tuning was accomplished with data collected in one afternoon. </i> </p>
<p>  <!-- 

<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/align_tshirt.gif"       height="185" width=""> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/color_stripe_front.gif" height="185" width=""> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/marker_marker.gif"      height="185" width=""> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/move_plate.gif"         height="185" width=""> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/socks_right.gif"        height="185" width=""> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/towel_stack.gif"        height="185" width=""> <br /> 
<i> Here you can see examples of visual foresight fine-tuned to perform basic control tasks in three entirely different environments. For the experiments, the target robot and environment was subtracted from RoboNet during pre-training. Fine-tuning was accomplished with data collected in one afternoon. </i> </p>

 -->  </p>
<p>We can now numerically evaluate if our pre-train controllers can pick up skills in new environments faster than a randomly initialized one. In each environment, we use a standard set of benchmark tasks to compare the performance of our pre-trained controller against the performance of a model trained only on data from the new environment. The results show that the fine-tuned model is ~4x more likely to complete the benchmark task than the one trained without RoboNet. Impressively, the pre-trained models can even slightly outperform models trained from scratch on significantly (5-20x) more data from the test environment. This suggests that transfer from RoboNet does indeed offer large performance gains compared to training from scratch!</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/graphs_franka_kuka.png" width="" /> <br /> <br />
<i> We compare the performance of fine-tuned models against their counterparts trained from scratch in two different test environments (with different robot platforms). </i> </p>
<p>Clearly fine-tuning is better than training from scratch, but is training on all of RoboNet always the best way to go? To test this, we compare pre-training on various subsets of RoboNet versus training from scratch. As seen before, the model pre-trained on all of RoboNet (excluding the Baxter platform) performs substantially better than the random initialization model. However, the “RoboNet pre-trained” model is outperformed by a model trained on a subset of RoboNet data collected on the Sawyer robot &#8211; the single-arm variant of Baxter.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/graphs_baxter.png" height="400" width="" /> <br /> <br />
<i> Models pre-trained on various subsets of RoboNet are compared to one trained from scratch in an unseen (during pre-training) Baxter control environment </i> </p>
<p>The similarities between the Baxter and Sawyer likely partly explain our results, but why does simply adding data to the training set hurt performance after fine-tuning? We theorize that this effect occurs due to model under-fitting. In other words, RoboNet is an extremely challenging dataset for a visual dynamics model, and imperfections in the model predictions result in bad control performance. However, larger models with more parameters tend to be more powerful, and thus make better predictions on RoboNet (visualized below). Note that increasing the number of parameters greatly improves prediction quality, but even large models with 500M parameters (middle column in the videos below) are still quite blurry. This suggests ample room for improvement, and we hope that the development of newer more powerful models will translate to better control performance in the future.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/robonet/compar_ppt.gif" height="450" width="" /> <br /> <br />
<i> We compare video prediction models of various size trained on RoboNet. A 75M parameter model (right-most column) generates significantly blurrier predictions than a large model with 500M parameters (center column). </i> </p>
<h1 id="final-thoughts">Final Thoughts</h1>
<p>This work takes the first step towards creating learned robotic agents that can operate in a wide range of environments and across different hardware. While our experiments primarily explore model-based reinforcement learning, we hope that RoboNet will inspire the broader robotics and reinforcement learning communities to investigate how to scale model-based <em>or</em> model-free RL algorithms to meet the complexity and diversity of the real world.</p>
<p>Since the dataset is extensible, we encourage other researchers to <a href="https://docs.google.com/forms/d/e/1FAIpQLSeV1XGvPQ6xTyEKGoTIbJWbKOCsUJ4gTRJ5fOQMWmlBowQwQQ/viewform" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">contribute</a> the data generated from their experiments back into RoboNet. After all, any data containing robot telemetry and video could be useful to someone else, so long as it contains the right documentation. In the long term, we believe this process will iteratively strengthen the dataset, and thus allow our algorithms that use  it to achieve greater levels of generalization across tasks, environments, robots, and experimental set-ups.</p>
<p>For more information please refer to the the <a href="https://www.robonet.wiki/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">project website</a>. We’ve also open sourced our <a href="https://github.com/SudeepDasari/RoboNet" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">code-base</a> and the entire <a href="https://github.com/SudeepDasari/RoboNet/wiki/Getting-Started" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RoboNet dataset</a>.</p>
<p>Finally, I would like to thank Sergey Levine, Chelsea Finn, and Frederik Ebert for their helpful feedback on this post.</p>
<p>  This article was initially published on the BAIR blog, and appears here with the authors’ permission.</p>
<p>This blog post was based on the following paper:</p>
<ul>
<li><strong><a href="https://arxiv.org/abs/1910.11215" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RoboNet: Large-Scale Multi-Robot Learning</a></strong>.<br /> S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, C. Finn.<br /> In Conference on Robot Learning, 2019.</li>
</ul>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Look then listen: Pre-learning environment representations for data-efficient neural instruction following</title>
		<link>https://robohub.org/look-then-listen-pre-learning-environment-representations-for-data-efficient-neural-instruction-following/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 06 Nov 2019 02:37:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/look-then-listen-pre-learning-environment-representations-for-data-efficient-neural-instruction-following/</guid>

					<description><![CDATA[When learning to follow natural language instructions, neural networks tend to
be very data hungry – they require a huge number of examples pairing language
with actions in order to learn effectively.  This post is about reducing those
heavy data requi...]]></description>
										<content:encoded><![CDATA[<p><strong>By David Gaddy</strong></p>
<p>When learning to follow natural language instructions, neural networks tend to be very data hungry – they require a huge number of examples pairing language with actions in order to learn effectively.  This post is about reducing those heavy data requirements by first watching actions in the environment before moving on to learning from language data.  Inspired by the idea that it is easier to map language to meanings that have already been formed, we introduce a semi-supervised approach that aims to separate the formation of abstractions from the learning of language.  <span id="more-151932"></span></p>
<p>Empirically, we find that pre-learning of patterns in the environment can help us learn grounded language with much less data.</p>
<p>Before we dive into the details, let’s look at an example to see why neural networks struggle to learn from smaller amounts of data.  For now, we’ll use examples from the <a href="https://shrdlurn.sidaw.xyz/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SHRDLURN block stacking task</a>, but later we’ll look at results on another environment.</p>
<p>Let’s put ourselves in the shoes of a model that is learning to follow instructions.  Suppose we are given the single training example below, which pairs a language command with an action in the environment:</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/look-then-listen/training_example.PNG" />  </p>
<p>This example tells us that if we are in state (a) and are trying to follow the instruction (b), the correct output for our model is the state (c).  Before learning, the model doesn’t know anything about language, so we must rely on examples like the one shown to figure out the meaning of the words.  After learning, we will be given new environment states and new instructions, and the model’s job is to choose the correct output states from executing the instructions.  First let’s consider a simple case where we get the exact same language, but the environment state is different, like the one shown here:</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/look-then-listen/new_state.PNG" />  </p>
<p>On this new state, the model has many different possible outputs that it could consider.  Here are just a few:</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/look-then-listen/many_options.PNG" />  </p>
<p>Some of these outputs seem reasonable to a human, like stacking red blocks on orange blocks or stacking red blocks on the left, but others are kind of strange, like generating a completely unrelated configuration of blocks.  To a neural network with no prior knowledge, however, all of these options look plausible.</p>
<p>A human learning a new language might approach this task by reasoning about possible meanings of the language that are consistent with the given example and choosing states that correspond to those meanings.  The set of possible meanings to consider comes from prior knowledge about what types of things might happen in an environment and how we can talk about them.  In this context, a meaning is an abstract transformation that we can apply to states to get new states.  For example, if someone saw the training instance above paired with language they didn’t understand, they might focus on two possible meanings for the instruction: it could be telling us to stack red blocks on orange blocks, or it could be telling us to stack a red block on the leftmost position.</p>
<p style="text-align:center;"> <img decoding="async" width="750" src="https://bair.berkeley.edu/static/blog/look-then-listen/limited_options.PNG" />  </p>
<p>Although we don’t know which of these two options is correct – both are plausible given the evidence – we now have many fewer options and might easily distinguish between them with just one or two more related examples.  Having a set of pre-formed meanings makes learning easier because the meanings constrain the space of possible outputs that must be considered.</p>
<p>In fact, pre-formed meanings do even more than just restricting the number of choices, because once we have chosen a meaning to pair with the language, it specifies the correct way to generalize across a wide variety of different initial environment states.  For example, consider the following transitions:</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/look-then-listen/generalize.PNG" />  </p>
<p>If we know in advance that all of these transitions belong together in a single semantic group (adding a red block on the left), learning language becomes easier because we can map to the group instead of the individual transitions. An end-to-end network that doesn’t start with any grouping of transitions has a much harder time because it has to learn the correct way to generalize across initial states.  One approach used by a long line of past work has been to provide the learner with a manually defined set of abstractions called logical forms.  In contrast, we take a more data-driven approach where we learn abstractions from unsupervised (language-free) data instead.</p>
<p>In this work, we help a neural network learn language with fewer examples by first learning abstractions from language-free observations of actions in an environment.  The idea here is that if the model sees lots of actions happening in an environment, perhaps it can pick up on patterns in what tends to be done, and these patterns might give hints at what abstractions are useful.  Our pre-learned abstractions can make language learning easier by constraining the space of outputs we need to consider and guiding generalization across different environment states.</p>
<p>We break up learning into two phases: an environment learning phase where our agent builds abstractions from language-free observation of the environment, and a language learning phase where natural language instructions are mapped to the pre-learned abstractions.  The motivation for this setup is that language-free observations of the environment are often easier to get than interactions paired with language, so we should use the cheaper unlabeled data to help us learn with less language data.  For example, a virtual assistant could learn with data from regular smartphone use, or in the longer term robots might be able to learn by watching humans naturally interact with the world. In the environments we are using in this post, we don’t have a natural source of unlabeled observations, so we generate the environment data synthetically.</p>
<h1 id="method">Method</h1>
<p>Now we’re ready to dive into our method.  We’ll start with the environment learning phase, where we will learn abstractions by observing an agent, such as a human, acting in the environment.  Our approach during this phase will be to create a type of autoencoder of the state transitions (actions) that we see, shown below:</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/look-then-listen/environment_learning_serif.PNG" />  </p>
<p>The encoder takes in the states before and after the transition and computes a representation of the transition itself.  The decoder takes that transition representation from the encoder and must use it to recreate the final state from the initial one.  The encoder and decoder architectures will be task specific, but use generic components such as convolutions or LSTMs.  For example, in the block stacking task states are represented as a grid and we use a convolutional architecture.  We train using a standard cross-entropy loss on the decoder’s output state, and after training we will use the representation passed between the encoder and decoder as our learned abstraction.</p>
<p>One thing that this autoencoder will learn is which type of transitions tend to happen, because the model will learn to only output transitions like the ones it sees during training.  In addition, this model will learn to <em>group</em> different transitions.  This grouping happens because the representation between the encoder and decoder acts as an information bottleneck, and its limited capacity forces the model to reuse the same representation vector for multiple different transitions.  We find that often the groupings it chooses tend to be semantically meaningful because representations that align with the semantics of the environment tend to be the most compact.</p>
<p>After environment learning pre-training, we are ready to move on to learning language.  For the language learning phase, we will start with the decoder that we pre-trained during environment learning (“action decoder” in the figures above and below).  The decoder maps from our learned representation space to particular state outputs.  To learn language, we now just need to introduce a language encoder module that maps from language into the representation space and train it by backpropagating through the decoder.  The model structure is shown in the figure below.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/look-then-listen/language_learning_serif.PNG" />  </p>
<p>The model in this phase looks a lot like other encoder-decoder models used previously for instruction following tasks, but now the pre-trained decoder can constrain the output and help control generalization.</p>
<h1 id="results">Results</h1>
<p>Now let’s look at some results.  We’ll compare our method to an end-to-end neural model, which has an identical neural architecture to our ultimate language learning model but without any environment learning pre-training of the decoder.  First we test on the <a href="https://shrdlurn.sidaw.xyz/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SHURDLURN block stacking task</a>, a task that is especially challenging for neural models because it requires learning with just tens of examples.  A baseline neural model gets an accuracy of 18% on the task, but with our environment learning pre-training, the model reaches 28%, an improvement of ten absolute percentage points.</p>
<p>We also tested our method on a string manipulation task where we learn to execute instructions like “insert the letters vw after every vowel” on a string of characters.  The chart below shows accuracy as we vary the amount of data for both the baseline end-to-end model and the model with our pre-training procedure.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/look-then-listen/string_chart_serif.PNG" />  </p>
<p>As shown above, using our pre-training method leads to much more data-efficient language learning compared to learning from scratch.  By pre-learning abstractions from the environment, our method increases data efficiency by more than an order of magnitude.  To learn more about our method, including some additional performance-improving tricks and an analysis of what pre-training learns, check out our paper from ACL 2019: <a href="https://arxiv.org/abs/1907.09671" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">https://arxiv.org/abs/1907.09671</a>.</p>
<p>This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Functional RL with Keras and Tensorflow Eager</title>
		<link>https://robohub.org/functional-rl-with-keras-and-tensorflow-eager/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sun, 20 Oct 2019 23:05:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/functional-rl-with-keras-and-tensorflow-eager/</guid>

					<description><![CDATA[<p>In this blog post, we explore a functional paradigm for implementing
<a href="https://en.wikipedia.org/wiki/Reinforcement_learning" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reinforcement learning</a>
(RL) algorithms. The paradigm will be that developers write the numerics of
their algorithm as independent, pure functions, and then use a library to
compile them into <em>policies</em> that can be trained at scale. We share how these
ideas were implemented in <a href="https://rllib.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RLlib</a>&#8217;s <a href="https://ray.readthedocs.io/en/latest/rllib-concepts.html#building-policies-in-tensorflow" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">policy builder
API</a>,
eliminating thousands of lines of &#8220;glue&#8221; code and bringing support for
<a href="https://ray.readthedocs.io/en/latest/rllib-models.html#tensorflow-models" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Keras</a>
and <a href="https://ray.readthedocs.io/en/latest/rllib-concepts.html#building-policies-in-tensorflow-eager" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">TensorFlow
2.0</a>.</p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/1*1EwDu6skRrPkNPx_fzpVbg.png">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img01.png"><br></p>

<!--more-->

<h3>Why Functional Programming?</h3>

<p>One of the key ideas behind functional programming is that programs can be
composed largely of pure functions, i.e., functions whose outputs are entirely
determined by their inputs. Here less is more: by imposing restrictions on what
functions can do, we gain the ability to more easily reason about and
manipulate their execution.</p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/0*D11A8FaF53k76olV">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img02.png"><br></p>

<p>In TensorFlow, such functions of tensors can be executed either
<strong>symbolically</strong> with placeholder inputs or <strong>eagerly</strong> with real tensor
values. Since such functions have no side-effects, they have the same effect on
inputs whether they are called once symbolically or many times eagerly.</p>

<h3>Functional Reinforcement Learning</h3>

<p>Consider the following loss function over agent rollout data, with current
state $s$, actions $a$, returns $r$, and policy $\pi$:</p>

<p>If you&#8217;re not familiar with RL, all this function is saying is that we should
try to <em>improve the probability of good actions</em> (i.e., actions that increase
the future returns). Such a loss is at the core of <a href="https://spinningup.openai.com/en/latest/spinningup/rl_intro3.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">policy
gradient</a>
algorithms. As we will see, defining the loss is almost all you need to start
training a RL policy in RLlib.</p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/0*C410k6WuEQY9ChMF">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img03.png"><br><i>
Given a set of rollouts, the <a href="https://spinningup.openai.com/en/latest/spinningup/rl_intro3.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">policy gradient</a>
loss seeks to improve the probability of good actions (i.e., those that lead to
a win in this Pong example&#160;above).
</i>
</p>

<p>A straightforward translation into Python is as follows. Here, the loss
function takes $(\pi, s, a, r)$, computes $\pi(s, a)$ as a discrete action
distribution, and returns the log probability of the actions multiplied by the
returns:</p>

<div><div><pre><code><span>def</span> <span>loss</span><span>(</span><span>model</span><span>,</span> <span>s</span><span>:</span> <span>Tensor</span><span>,</span> <span>a</span><span>:</span>  <span>Tensor</span><span>,</span> <span>r</span><span>:</span> <span>Tensor</span><span>)</span> <span>-&#62;</span> <span>Tensor</span><span>:</span>
    <span>logits</span> <span>=</span> <span>model</span><span>.</span><span>forward</span><span>(</span><span>s</span><span>)</span>
    <span>action_dist</span> <span>=</span> <span>Categorical</span><span>(</span><span>logits</span><span>)</span>
    <span>return</span> <span>-</span><span>tf</span><span>.</span><span>reduce_mean</span><span>(</span><span>action_dist</span><span>.</span><span>logp</span><span>(</span><span>a</span><span>)</span> <span>*</span> <span>r</span><span>)</span>
</code></pre></div></div>

<p>There are multiple benefits to this functional definition. First, notice that
loss reads quite naturally&#8202;&#8212;&#8202;<strong>there are no placeholders, control loops, access
of external variables, or class members</strong> as commonly seen in RL
implementations. Second, since it doesn&#8217;t mutate external state, it is
compatible with both TF graph and eager mode execution.</p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/0*GTmsc5AQZbE4f9qY">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img04.png"><br><i>
In contrast to a class-based API, in which class methods can access arbitrary
parts of the class state, a functional API builds policies from loosely coupled
pure functions.
</i>
</p>

<p>In this blog we explore defining RL algorithms as collections of such pure
functions. The paradigm will be that developers write the numerics of their
algorithm as independent, pure functions, and then use a RLlib helper function
to compile them into <em>policies</em> that can be trained at scale. This proposal is
implemented concretely in the RLlib library.</p>

<h3>Functional RL with&#160;RLlib</h3>

<p><a href="https://ray.readthedocs.io/en/latest/rllib.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RLlib</a> is an open-source
library for reinforcement learning that offers both high scalability and a
unified API for a variety of applications. It offers a <a href="https://ray.readthedocs.io/en/latest/rllib-algorithms.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">wide range of scalable
RL algorithms</a>.</p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/0*2wpxKQ_TBBQW7Lhe">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img05.png"><br><i>
Example of how RLlib scales algorithms, in this case with distributed
synchronous sampling.
</i>
</p>

<p>Given the increasing popularity of PyTorch (i.e., imperative execution) and the
imminent release of TensorFlow 2.0, we saw the opportunity to improve RLlib&#8217;s
developer experience with a functional rewrite of RLlib&#8217;s algorithms. The major
goals were to:</p>

<p><strong>Improve the RL debugging experience</strong></p>

<ul><li>Allow eager execution to be used for any algorithm with just an&#8202;&#8212;&#8202;eager
flag, enabling easy <code>print()</code> debugging.</li>
</ul><p><strong>Simplify new algorithm development</strong></p>

<ul><li>Make algorithms easier to customize and understand by replacing monolithic
&#8220;Agent&#8221; classes with policies built from collections of pure functions
(e.g., primitives provided by <a href="https://github.com/deepmind/trfl" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">TRFL</a>).</li>
  <li>Remove the need to manually declare tensor placeholders for TF.</li>
  <li>Unify the way TF and PyTorch policies are defined.</li>
</ul><h3>Policy Builder&#160;API</h3>

<p>The RLlib policy builder API for functional RL (stable in RLlib 0.7.4) involves
just two key functions:</p>

<ul><li><a href="https://ray.readthedocs.io/en/latest/rllib-concepts.html#building-policies-in-tensorflow" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">build_tf_policy</a>()</li>
  <li><a href="https://ray.readthedocs.io/en/latest/rllib-concepts.html#building-policies-in-pytorch" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">build_torch_policy</a>()</li>
</ul><p>At a high level, these builders take a number of <strong>function objects</strong> as input,
including a <code>loss_fn</code> similar to what you saw earlier, a <code>model_fn</code> to return a
neural network model given the algorithm config, and an <code>action_fn</code> to generate
action samples given model outputs. The actual API takes quite a few more
arguments, but these are the main ones. The builder compiles these functions
into a
<a href="https://ray.readthedocs.io/en/latest/rllib-concepts.html#policies" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">policy</a>
that can be queried for actions and improved over time given experiences:</p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/0*IFYYtLJg-FyGEI77">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img06.png"><br></p>

<p>These policies can be leveraged for single-agent, vector, and multi-agent
training in RLlib, which calls on them to determine how to interact with
environments:</p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/0*nVuy28pbOkgNygaD">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img07.png"><br></p>

<p>We&#8217;ve found the policy builder pattern general enough to port almost all of RLlib&#8217;s reference algorithms, including <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/a3c/a3c_tf_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">A2C</a>, <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/ppo/appo_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">APPO</a>, <a href="https://github.com/ray-project/ray/blob/b520f6141ecdd54496b0c26106f3df4442a5f91e/rllib/agents/ddpg/ddpg_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DDPG</a>, <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/dqn/dqn_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DQN</a>, <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/pg/pg_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PG</a>, <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/ppo/appo_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PPO</a>, <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/sac/sac_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SAC</a>, and <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/impala/vtrace_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">IMPALA</a> in TensorFlow, and <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/pg/torch_pg_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PG</a> / <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/a3c/a3c_torch_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">A2C</a> in PyTorch. While code readability is somewhat subjective, users have reported that the builder pattern makes it much easier to customize algorithms, especially in environments such as Jupyter notebooks. In addition, these refactorings have reduced the size of the algorithms by up to hundreds of lines of code <em>each</em>.</p>

<h3>Vanilla Policy Gradients Example</h3>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/0*I7SuZh-u1rl3Emfb">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img08.png"><br><i>
Visualization of the vanilla policy gradient loss function in&#160;RLlib.
</i>
</p>

<p>Let&#8217;s take a look at how the earlier loss example can be implemented concretely
using the builder pattern. We define <code>policy_gradient_loss</code>, which requires a
couple of tweaks for generality: (1) RLlib supplies the proper
<code>distribution_class</code> so the algorithm can work with any type of action space
(e.g., continuous or categorical), and (2) the experience data is held in a
<code>train_batch</code> dict that contains state, action, etc. tensors:</p>

<div><div><pre><code><span>def</span> <span>policy_gradient_loss</span><span>(</span>
        <span>policy</span><span>,</span> <span>model</span><span>,</span> <span>distribution_cls</span><span>,</span> <span>train_batch</span><span>):</span>
    <span>logits</span><span>,</span> <span>_</span> <span>=</span> <span>model</span><span>.</span><span>from_batch</span><span>(</span><span>train_batch</span><span>)</span>
    <span>action_dist</span> <span>=</span> <span>distribution_cls</span><span>(</span><span>logits</span><span>,</span> <span>model</span><span>)</span>
    <span>return</span> <span>-</span><span>tf</span><span>.</span><span>reduce_mean</span><span>(</span>
        <span>action_dist</span><span>.</span><span>logp</span><span>(</span><span>train_batch</span><span>[</span><span>&#8220;</span><span>actions</span><span>&#8221;</span><span>])</span> <span>*</span>
        <span>train_batch</span><span>[</span><span>&#8220;</span><span>returns</span><span>&#8221;</span><span>])</span>
</code></pre></div></div>

<p>To add the &#8220;returns&#8221; array to the batch, we need to define a postprocessing
function that calculates it as the <a href="https://spinningup.openai.com/en/latest/spinningup/rl_intro.html#reward-and-return" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">temporally discounted
reward</a>
over the trajectory:</p>

<p>We set $\gamma = 0.99$ when computing $R(T)$ below in code:</p>

<div><div><pre><code><span>from</span> <span>ray.rllib.evaluation.postprocessing</span> <span>import</span> <span>discount</span>

<span># Run for each trajectory collected from the environment</span>
<span>def</span> <span>calculate_returns</span><span>(</span><span>policy</span><span>,</span>
                      <span>batch</span><span>,</span>
                      <span>other_agent_batches</span><span>=</span><span>None</span><span>,</span>
                      <span>episode</span><span>=</span><span>None</span><span>):</span>
   <span>batch</span><span>[</span><span>&#8220;</span><span>returns</span><span>&#8221;</span><span>]</span> <span>=</span> <span>discount</span><span>(</span><span>batch</span><span>[</span><span>&#8220;</span><span>rewards</span><span>&#8221;</span><span>],</span> <span>0.99</span><span>)</span>
   <span>return</span> <span>batch</span>
</code></pre></div></div>

<p>Given these functions, we can then build the RLlib policy and
<a href="https://ray.readthedocs.io/en/latest/rllib-concepts.html#trainers" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">trainer</a>
(which coordinates the overall training workflow). The model and action
distribution are automatically supplied by RLlib if not specified:</p>

<div><div><pre><code><span>MyTFPolicy</span> <span>=</span> <span>build_tf_policy</span><span>(</span>
   <span>name</span><span>=</span><span>"MyTFPolicy"</span><span>,</span>
   <span>loss_fn</span><span>=</span><span>policy_gradient_loss</span><span>,</span>
   <span>postprocess_fn</span><span>=</span><span>calculate_returns</span><span>)</span>

<span>MyTrainer</span> <span>=</span> <span>build_trainer</span><span>(</span>
   <span>name</span><span>=</span><span>"MyCustomTrainer"</span><span>,</span> <span>default_policy</span><span>=</span><span>MyTFPolicy</span><span>)</span>
</code></pre></div></div>

<p>Now we can run this at the desired scale using
<a href="https://ray.readthedocs.io/en/latest/tune.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Tune</a>, in this example showing
a configuration using 128 CPUs and 1 GPU in a cluster:</p>

<div><div><pre><code><span>tune</span><span>.</span><span>run</span><span>(</span><span>MyTrainer</span><span>,</span>
    <span>config</span><span>=</span><span>{</span><span>&#8220;</span><span>env</span><span>&#8221;</span><span>:</span> <span>&#8220;</span><span>CartPole</span><span>-</span><span>v0</span><span>&#8221;</span><span>,</span>
            <span>&#8220;</span><span>num_workers</span><span>&#8221;</span><span>:</span> <span>128</span><span>,</span>
            <span>&#8220;</span><span>num_gpus</span><span>&#8221;</span><span>:</span> <span>1</span><span>})</span>
</code></pre></div></div>

<p>While this example <a href="https://github.com/ray-project/ray/blob/master/rllib/examples/custom_tf_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">(runnable
code)</a>
is only a basic algorithm, it demonstrates how a functional API can be concise,
readable, and highly scalable. When compared against the previous way to define
policies in RLlib using TF placeholders, <strong>the</strong> <strong>functional API uses ~3x
fewer lines of code (23 vs 81 lines),</strong> and also works in eager:</p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/1*hrzYi0u3I6uARF-XvfcLTQ.png">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img09.png"><br><i>
Comparing the <a href="https://github.com/ray-project/ray/blob/75ac016e2bd39060d14a302292546d9dbc49f6a2/python/ray/rllib/agents/pg/pg_policy_graph.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">legacy class-based API</a>
with the new <a href="https://github.com/ray-project/ray/blob/b520f6141ecdd54496b0c26106f3df4442a5f91e/rllib/agents/pg/pg_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">functional policy builder API</a>
Both policies implement the same behaviour, but the functional definition is
much&#160;shorter.
</i>
</p>

<h3>How the Policy Builder&#160;works</h3>

<p>Under the hood, <code>build_tf_policy</code> takes the supplied building blocks
(<code>model_fn</code>, <code>action_fn</code>, <code>loss_fn</code>, etc.) and compiles them into either a
<a href="https://github.com/ray-project/ray/blob/03a1b758526b2699a21e44a932bb2abdfe636f2b/rllib/policy/dynamic_tf_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DynamicTFPolicy</a>
or
<a href="https://github.com/ray-project/ray/blob/03a1b758526b2699a21e44a932bb2abdfe636f2b/rllib/policy/eager_tf_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">EagerTFPolicy</a>,
depending on if TF eager execution is enabled. The former implements graph-mode
execution (auto-defining placeholders dynamically), the latter eager execution.</p>

<p>The main difference between <code>DynamicTFPolicy</code> and <code>EagerTFPolicy</code> is how many
times they call the functions passed in. In either case, a <code>model_fn</code> is
invoked once to create a Model class. However, functions that involve tensor
operations are either called once in graph mode to build a symbolic computation
graph, or multiple times in eager mode on actual tensors. In the following
figures we show how these operations work together in blue and orange:</p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/1*4RKNH6Zt4-P82ZeoETKVog.png">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img10.png"><br></p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/0*1vTJnByqwRvPoyBV">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img11.png"><br><i>
Overview of a generated EagerTFPolicy. The policy passes the environment state
through model.forward(), which emits output logits. The model output
parameterizes a probability distribution over actions (&#8220;ActionDistribution&#8221;),
which can be used when sampling actions or training. The loss function operates
over batches of experiences. The model can provide additional methods such as a
value function (light orange) or other methods for computing Q values, etc.
(not shown) as needed by the loss function.
</i>
</p>

<p>This policy object is all RLlib needs to launch and scale RL training.
Intuitively, this is because it encapsulates how to compute actions and improve
the policy. External state such as that of the environment and RNN hidden state
is managed externally by RLlib, and does not need to be part of the policy
definition. The policy object is used in one of two ways depending on whether
we are computing rollouts or trying to improve the policy given a batch of
rollout data:</p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/0*-35jGBA7Gha9WnOA">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img12.gif"><br><i>
<b>Inference:</b> Forward pass to compute a single action. This only involves
querying the model, generating an action distribution, and sampling an action
from that distribution. In eager mode, this involves calling action_fn
<a href="https://github.com/ray-project/ray/blob/03a1b758526b2699a21e44a932bb2abdfe636f2b/rllib/agents/dqn/simple_q_policy.py#L111" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">
DQN example of an action sampler</a>,
which creates an action distribution / action sampler as relevant that is then
sampled&#160;from.
</i>
</p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/0*w4cJ0KPTM8QPX5Ex">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img13.gif"><br><i>
<b>Training:</b> Forward and backward pass to learn on a batch of experiences.
In this mode, we call the loss function to generate a scalar output which can
be used to optimize the model variables via SGD. In eager mode, both action_fn
and loss_fn are called to generate the action distribution and policy loss
respectively. Note that here we don&#8217;t show differentiation through action_fn,
but this does happen in algorithms such as&#160;DQN.
</i>
</p>

<h3>Loose Ends: State Management</h3>

<p>RL training inherently involves a lot of state. If algorithms are defined using
pure functions, where is the state held? In most cases it can be managed
automatically by the framework. There are three types of state that need to be
managed in RLlib:</p>

<ol><li><strong>Environment state</strong>: this includes the current state of the environment
and any recurrent state passed between policy steps. RLlib manages this
internally in its <a href="https://github.com/ray-project/ray/blob/master/rllib/evaluation/rollout_worker.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">rollout
worker</a>
implementation.</li>
  <li><strong>Model state</strong>: these are the policy parameters we are trying to learn via
an RL loss. These variables must be accessible and optimized in the same way
for both graph and eager mode. Fortunately,
<a href="https://www.tensorflow.org/guide/keras" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Keras</a> models can be used in either
mode. RLlib provides a <a href="https://ray.readthedocs.io/en/latest/rllib-models.html#tensorflow-models" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">customizable model class
(TFModelV2)</a>
based on the object-oriented Keras style to hold policy parameters.</li>
  <li><strong>Training workflow state</strong>: state for managing training, e.g., the
annealing schedule for various hyperparameters, steps since last update, and so
on. RLlib lets algorithm authors add <a href="https://github.com/ray-project/ray/blob/b520f6141ecdd54496b0c26106f3df4442a5f91e/rllib/agents/ppo/ppo_policy.py#L284" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">mixin
classes</a>
to policies that can hold any such extra variables.</li>
</ol><h3>Loose ends: Eager&#160;Overhead</h3>

<p>Next we investigate RLlib&#8217;s eager mode performance with <a href="https://www.tensorflow.org/beta/tutorials/eager/tf_function" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">eager
tracing</a> on or
off. As shown in the below figure, tracing greatly improves performance.
However, the tradeoff is that Python operations such as print may not be called
each time. For this reason, tracing is off by default in RLlib, but can be
enabled with &#8220;eager_tracing&#8221;: True. In addition, you can also set
&#8220;no_eager_on_workers&#8221; to enable eager only for learning but disable it for
inference:</p>

<p>
&#60;!--
<img src="https://cdn-images-1.medium.com/max/800/0*YvplJKQocQXjclg5">
--&#62;
<img src="https://bair.berkeley.edu/static/blog/functional/img14.png"><br></p>

<p>Eager inference and gradient overheads measured using <code>rllib train --run=PG
--env=&#60;env&#62; [ --eager [ --trace]]</code> on a laptop processor. With tracing off, eager
imposes a significant overhead for small batch operations. However it is often
as fast or faster than graph mode when tracing is&#160;enabled.</p>

<h3>Conclusion</h3>

<p>To recap, in this blog post we propose using ideas from functional programming
to simplify the development of RL algorithms. We implement and validate these
ideas in RLlib. Beyond making it easy to support new features such as eager
execution, we also find the functional paradigm leads to substantially more
concise and understandable code. Try it out yourself with <code>pip install
ray[rllib]</code> or by checking out the
<a href="https://ray.readthedocs.io/en/latest/rllib.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">docs</a> and <a href="https://github.com/ray-project/ray/tree/master/rllib" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">source
code</a>.</p>

<p>If you&#8217;re interested in helping improve RLlib, we&#8217;re also <a href="https://jobs.lever.co/anyscale" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">hiring</a>.</p>]]></description>
										<content:encoded><![CDATA[<p><strong>By Eric Liang and Richard Liaw and Clement Gehring</strong></p>
<p>In this blog post, we explore a functional paradigm for implementing <a href="https://en.wikipedia.org/wiki/Reinforcement_learning" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reinforcement learning</a> (RL) algorithms. The paradigm will be that developers write the numerics of their algorithm as independent, pure functions, and then use a library to compile them into <em>policies</em> that can be trained at scale. We share how these ideas were implemented in <a href="https://rllib.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RLlib</a>’s <a href="https://ray.readthedocs.io/en/latest/rllib-concepts.html#building-policies-in-tensorflow" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">policy builder API</a>, eliminating thousands of lines of “glue” code and bringing support for <a href="https://ray.readthedocs.io/en/latest/rllib-models.html#tensorflow-models" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Keras</a> and <a href="https://ray.readthedocs.io/en/latest/rllib-concepts.html#building-policies-in-tensorflow-eager" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">TensorFlow 2.0</a>.</p>
<p style="text-align:center;"> </p>
<p>  <span id="more-149778"></span>  </p>
<img decoding="async" src="https://robohub.org/wp-content/uploads/2019/10/RlibTensorflow.png" alt="" width="900" height="364" class="aligncenter size-full wp-image-150897" srcset="https://robohub.org/wp-content/uploads/2019/10/RlibTensorflow.png 900w, https://robohub.org/wp-content/uploads/2019/10/RlibTensorflow-425x172.png 425w, https://robohub.org/wp-content/uploads/2019/10/RlibTensorflow-768x311.png 768w" sizes="(max-width: 900px) 100vw, 900px" />
<h3 id="why-functional-programming">Why Functional Programming?</h3>
<p>One of the key ideas behind functional programming is that programs can be composed largely of pure functions, i.e., functions whose outputs are entirely determined by their inputs. Here less is more: by imposing restrictions on what functions can do, we gain the ability to more easily reason about and manipulate their execution.</p>
<p style="text-align:center;"> <!-- <img decoding="async" src="https://cdn-images-1.medium.com/max/800/0*D11A8FaF53k76olV"> --> <img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img02.png" />  </p>
<p>In TensorFlow, such functions of tensors can be executed either <strong>symbolically</strong> with placeholder inputs or <strong>eagerly</strong> with real tensor values. Since such functions have no side-effects, they have the same effect on inputs whether they are called once symbolically or many times eagerly.</p>
<h3 id="functional-reinforcement-learning">Functional Reinforcement Learning</h3>
<p>Consider the following loss function over agent rollout data, with current state $s$, actions $a$, returns $r$, and policy $\pi$:</p>
<p>  <script type="math/tex; mode=display">L(s, a, r) = -[\log \pi(s, a)] \cdot r</script>  </p>
<p>If you’re not familiar with RL, all this function is saying is that we should try to <em>improve the probability of good actions</em> (i.e., actions that increase the future returns). Such a loss is at the core of <a href="https://spinningup.openai.com/en/latest/spinningup/rl_intro3.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">policy gradient</a> algorithms. As we will see, defining the loss is almost all you need to start training a RL policy in RLlib.</p>
<p style="text-align:center;"> <!-- <img decoding="async" src="https://cdn-images-1.medium.com/max/800/0*C410k6WuEQY9ChMF"> --> <img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img03.png" /> <br /> <i> Given a set of rollouts, the <a href="https://spinningup.openai.com/en/latest/spinningup/rl_intro3.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">policy gradient</a> loss seeks to improve the probability of good actions (i.e., those that lead to a win in this Pong example above). </i> </p>
<p>A straightforward translation into Python is as follows. Here, the loss function takes $(\pi, s, a, r)$, computes $\pi(s, a)$ as a discrete action distribution, and returns the log probability of the actions multiplied by the returns:</p>
<div class="language-python highlighter-rouge">
<div class="highlight">
<pre class="highlight"><code><span class="k">def</span> <span class="nf">loss</span><span class="p">(</span><span class="n">model</span><span class="p">,</span> <span class="n">s</span><span class="p">:</span> <span class="n">Tensor</span><span class="p">,</span> <span class="n">a</span><span class="p">:</span>  <span class="n">Tensor</span><span class="p">,</span> <span class="n">r</span><span class="p">:</span> <span class="n">Tensor</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="n">Tensor</span><span class="p">:</span>
    <span class="n">logits</span> <span class="o">=</span> <span class="n">model</span><span class="o">.</span><span class="n">forward</span><span class="p">(</span><span class="n">s</span><span class="p">)</span>
    <span class="n">action_dist</span> <span class="o">=</span> <span class="n">Categorical</span><span class="p">(</span><span class="n">logits</span><span class="p">)</span>
    <span class="k">return</span> <span class="o">-</span><span class="n">tf</span><span class="o">.</span><span class="n">reduce_mean</span><span class="p">(</span><span class="n">action_dist</span><span class="o">.</span><span class="n">logp</span><span class="p">(</span><span class="n">a</span><span class="p">)</span> <span class="o">*</span> <span class="n">r</span><span class="p">)</span>
</code></pre>
</div>
</div>
<p>There are multiple benefits to this functional definition. First, notice that loss reads quite naturally — <strong>there are no placeholders, control loops, access of external variables, or class members</strong> as commonly seen in RL implementations. Second, since it doesn’t mutate external state, it is compatible with both TF graph and eager mode execution.</p>
<p style="text-align:center;"> <!-- <img decoding="async" src="https://cdn-images-1.medium.com/max/800/0*GTmsc5AQZbE4f9qY"> --> <img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img04.png" /> <br /> <i> In contrast to a class-based API, in which class methods can access arbitrary parts of the class state, a functional API builds policies from loosely coupled pure functions. </i> </p>
<p>In this blog we explore defining RL algorithms as collections of such pure functions. The paradigm will be that developers write the numerics of their algorithm as independent, pure functions, and then use a RLlib helper function to compile them into <em>policies</em> that can be trained at scale. This proposal is implemented concretely in the RLlib library.</p>
<h3 id="functional-rl-withrllib">Functional RL with RLlib</h3>
<p><a href="https://ray.readthedocs.io/en/latest/rllib.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RLlib</a> is an open-source library for reinforcement learning that offers both high scalability and a unified API for a variety of applications. It offers a <a href="https://ray.readthedocs.io/en/latest/rllib-algorithms.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">wide range of scalable RL algorithms</a>.</p>
<p style="text-align:center;"> <!-- <img decoding="async" src="https://cdn-images-1.medium.com/max/800/0*2wpxKQ_TBBQW7Lhe"> --> <img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img05.png" /> <br /> <i> Example of how RLlib scales algorithms, in this case with distributed synchronous sampling. </i> </p>
<p>Given the increasing popularity of PyTorch (i.e., imperative execution) and the imminent release of TensorFlow 2.0, we saw the opportunity to improve RLlib’s developer experience with a functional rewrite of RLlib’s algorithms. The major goals were to:</p>
<p><strong>Improve the RL debugging experience</strong></p>
<ul>
<li>Allow eager execution to be used for any algorithm with just an — eager flag, enabling easy <code class="highlighter-rouge">print()</code> debugging.</li>
</ul>
<p><strong>Simplify new algorithm development</strong></p>
<ul>
<li>Make algorithms easier to customize and understand by replacing monolithic “Agent” classes with policies built from collections of pure functions (e.g., primitives provided by <a href="https://github.com/deepmind/trfl" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">TRFL</a>).</li>
<li>Remove the need to manually declare tensor placeholders for TF.</li>
<li>Unify the way TF and PyTorch policies are defined.</li>
</ul>
<h3 id="policy-builderapi">Policy Builder API</h3>
<p>The RLlib policy builder API for functional RL (stable in RLlib 0.7.4) involves just two key functions:</p>
<ul>
<li><a href="https://ray.readthedocs.io/en/latest/rllib-concepts.html#building-policies-in-tensorflow" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">build_tf_policy</a>()</li>
<li><a href="https://ray.readthedocs.io/en/latest/rllib-concepts.html#building-policies-in-pytorch" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">build_torch_policy</a>()</li>
</ul>
<p>At a high level, these builders take a number of <strong>function objects</strong> as input, including a <code class="highlighter-rouge">loss_fn</code> similar to what you saw earlier, a <code class="highlighter-rouge">model_fn</code> to return a neural network model given the algorithm config, and an <code class="highlighter-rouge">action_fn</code> to generate action samples given model outputs. The actual API takes quite a few more arguments, but these are the main ones. The builder compiles these functions into a <a href="https://ray.readthedocs.io/en/latest/rllib-concepts.html#policies" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">policy</a> that can be queried for actions and improved over time given experiences:</p>
<p style="text-align:center;"> <!-- <img decoding="async" src="https://cdn-images-1.medium.com/max/800/0*IFYYtLJg-FyGEI77"> --> <img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img06.png" />  </p>
<p>These policies can be leveraged for single-agent, vector, and multi-agent training in RLlib, which calls on them to determine how to interact with environments:</p>
<p style="text-align:center;"> <!-- <img decoding="async" src="https://cdn-images-1.medium.com/max/800/0*nVuy28pbOkgNygaD"> --> <img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img07.png" />  </p>
<p>We’ve found the policy builder pattern general enough to port almost all of RLlib’s reference algorithms, including <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/a3c/a3c_tf_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">A2C</a>, <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/ppo/appo_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">APPO</a>, <a href="https://github.com/ray-project/ray/blob/b520f6141ecdd54496b0c26106f3df4442a5f91e/rllib/agents/ddpg/ddpg_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DDPG</a>, <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/dqn/dqn_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DQN</a>, <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/pg/pg_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PG</a>, <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/ppo/appo_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PPO</a>, <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/sac/sac_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SAC</a>, and <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/impala/vtrace_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">IMPALA</a> in TensorFlow, and <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/pg/torch_pg_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PG</a> / <a href="https://github.com/ray-project/ray/blob/master/rllib/agents/a3c/a3c_torch_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">A2C</a> in PyTorch. While code readability is somewhat subjective, users have reported that the builder pattern makes it much easier to customize algorithms, especially in environments such as Jupyter notebooks. In addition, these refactorings have reduced the size of the algorithms by up to hundreds of lines of code <em>each</em>.</p>
<h3 id="vanilla-policy-gradients-example">Vanilla Policy Gradients Example</h3>
<p style="text-align:center;"> <!-- <img decoding="async" src="https://cdn-images-1.medium.com/max/800/0*I7SuZh-u1rl3Emfb"> --> <img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img08.png" /> <br /> <i> Visualization of the vanilla policy gradient loss function in RLlib. </i> </p>
<p>Let’s take a look at how the earlier loss example can be implemented concretely using the builder pattern. We define <code class="highlighter-rouge">policy_gradient_loss</code>, which requires a couple of tweaks for generality: (1) RLlib supplies the proper <code class="highlighter-rouge">distribution_class</code> so the algorithm can work with any type of action space (e.g., continuous or categorical), and (2) the experience data is held in a <code class="highlighter-rouge">train_batch</code> dict that contains state, action, etc. tensors:</p>
<div class="language-python highlighter-rouge">
<div class="highlight">
<pre class="highlight"><code><span class="k">def</span> <span class="nf">policy_gradient_loss</span><span class="p">(</span>
        <span class="n">policy</span><span class="p">,</span> <span class="n">model</span><span class="p">,</span> <span class="n">distribution_cls</span><span class="p">,</span> <span class="n">train_batch</span><span class="p">):</span>
    <span class="n">logits</span><span class="p">,</span> <span class="n">_</span> <span class="o">=</span> <span class="n">model</span><span class="o">.</span><span class="n">from_batch</span><span class="p">(</span><span class="n">train_batch</span><span class="p">)</span>
    <span class="n">action_dist</span> <span class="o">=</span> <span class="n">distribution_cls</span><span class="p">(</span><span class="n">logits</span><span class="p">,</span> <span class="n">model</span><span class="p">)</span>
    <span class="k">return</span> <span class="o">-</span><span class="n">tf</span><span class="o">.</span><span class="n">reduce_mean</span><span class="p">(</span>
        <span class="n">action_dist</span><span class="o">.</span><span class="n">logp</span><span class="p">(</span><span class="n">train_batch</span><span class="p">[</span><span class="err">“</span><span class="n">actions</span><span class="err">”</span><span class="p">])</span> <span class="o">*</span>
        <span class="n">train_batch</span><span class="p">[</span><span class="err">“</span><span class="n">returns</span><span class="err">”</span><span class="p">])</span>
</code></pre>
</div>
</div>
<p>To add the “returns” array to the batch, we need to define a postprocessing function that calculates it as the <a href="https://spinningup.openai.com/en/latest/spinningup/rl_intro.html#reward-and-return" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">temporally discounted reward</a> over the trajectory:</p>
<p>  <script type="math/tex; mode=display">R(\tau) = \sum_{t=0}^{\infty}{\gamma^tr_t}</script>  </p>
<p>We set $\gamma = 0.99$ when computing $R(T)$ below in code:</p>
<div class="language-python highlighter-rouge">
<div class="highlight">
<pre class="highlight"><code><span class="kn">from</span> <span class="nn">ray.rllib.evaluation.postprocessing</span> <span class="kn">import</span> <span class="n">discount</span>

<span class="c"># Run for each trajectory collected from the environment</span>
<span class="k">def</span> <span class="nf">calculate_returns</span><span class="p">(</span><span class="n">policy</span><span class="p">,</span>
                      <span class="n">batch</span><span class="p">,</span>
                      <span class="n">other_agent_batches</span><span class="o">=</span><span class="bp">None</span><span class="p">,</span>
                      <span class="n">episode</span><span class="o">=</span><span class="bp">None</span><span class="p">):</span>
   <span class="n">batch</span><span class="p">[</span><span class="err">“</span><span class="n">returns</span><span class="err">”</span><span class="p">]</span> <span class="o">=</span> <span class="n">discount</span><span class="p">(</span><span class="n">batch</span><span class="p">[</span><span class="err">“</span><span class="n">rewards</span><span class="err">”</span><span class="p">],</span> <span class="mf">0.99</span><span class="p">)</span>
   <span class="k">return</span> <span class="n">batch</span>
</code></pre>
</div>
</div>
<p>Given these functions, we can then build the RLlib policy and <a href="https://ray.readthedocs.io/en/latest/rllib-concepts.html#trainers" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">trainer</a> (which coordinates the overall training workflow). The model and action distribution are automatically supplied by RLlib if not specified:</p>
<div class="language-python highlighter-rouge">
<div class="highlight">
<pre class="highlight"><code><span class="n">MyTFPolicy</span> <span class="o">=</span> <span class="n">build_tf_policy</span><span class="p">(</span>
   <span class="n">name</span><span class="o">=</span><span class="s">"MyTFPolicy"</span><span class="p">,</span>
   <span class="n">loss_fn</span><span class="o">=</span><span class="n">policy_gradient_loss</span><span class="p">,</span>
   <span class="n">postprocess_fn</span><span class="o">=</span><span class="n">calculate_returns</span><span class="p">)</span>

<span class="n">MyTrainer</span> <span class="o">=</span> <span class="n">build_trainer</span><span class="p">(</span>
   <span class="n">name</span><span class="o">=</span><span class="s">"MyCustomTrainer"</span><span class="p">,</span> <span class="n">default_policy</span><span class="o">=</span><span class="n">MyTFPolicy</span><span class="p">)</span>
</code></pre>
</div>
</div>
<p>Now we can run this at the desired scale using <a href="https://ray.readthedocs.io/en/latest/tune.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Tune</a>, in this example showing a configuration using 128 CPUs and 1 GPU in a cluster:</p>
<div class="language-python highlighter-rouge">
<div class="highlight">
<pre class="highlight"><code><span class="n">tune</span><span class="o">.</span><span class="n">run</span><span class="p">(</span><span class="n">MyTrainer</span><span class="p">,</span>
    <span class="n">config</span><span class="o">=</span><span class="p">{</span><span class="err">“</span><span class="n">env</span><span class="err">”</span><span class="p">:</span> <span class="err">“</span><span class="n">CartPole</span><span class="o">-</span><span class="n">v0</span><span class="err">”</span><span class="p">,</span>
            <span class="err">“</span><span class="n">num_workers</span><span class="err">”</span><span class="p">:</span> <span class="mi">128</span><span class="p">,</span>
            <span class="err">“</span><span class="n">num_gpus</span><span class="err">”</span><span class="p">:</span> <span class="mi">1</span><span class="p">})</span>
</code></pre>
</div>
</div>
<p>While this example <a href="https://github.com/ray-project/ray/blob/master/rllib/examples/custom_tf_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">(runnable code)</a> is only a basic algorithm, it demonstrates how a functional API can be concise, readable, and highly scalable. When compared against the previous way to define policies in RLlib using TF placeholders, <strong>the</strong> <strong>functional API uses ~3x fewer lines of code (23 vs 81 lines),</strong> and also works in eager:</p>
<p style="text-align:center;">
<!--
<img decoding="async" src="https://cdn-images-1.medium.com/max/800/1*hrzYi0u3I6uARF-XvfcLTQ.png">
--><br />
<img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img09.png" /><br />
<br />
<i><br />
Comparing the <a href="https://github.com/ray-project/ray/blob/75ac016e2bd39060d14a302292546d9dbc49f6a2/python/ray/rllib/agents/pg/pg_policy_graph.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">legacy class-based API</a><br />
with the new <a href="https://github.com/ray-project/ray/blob/b520f6141ecdd54496b0c26106f3df4442a5f91e/rllib/agents/pg/pg_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">functional policy builder API</a><br />
Both policies implement the same behaviour, but the functional definition is<br />
much shorter.<br />
</i>
</p>
<h3 id="how-the-policy-builderworks">How the Policy Builder works</h3>
<p>Under the hood, <code class="highlighter-rouge">build_tf_policy</code> takes the supplied building blocks (<code class="highlighter-rouge">model_fn</code>, <code class="highlighter-rouge">action_fn</code>, <code class="highlighter-rouge">loss_fn</code>, etc.) and compiles them into either a <a href="https://github.com/ray-project/ray/blob/03a1b758526b2699a21e44a932bb2abdfe636f2b/rllib/policy/dynamic_tf_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DynamicTFPolicy</a> or <a href="https://github.com/ray-project/ray/blob/03a1b758526b2699a21e44a932bb2abdfe636f2b/rllib/policy/eager_tf_policy.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">EagerTFPolicy</a>, depending on if TF eager execution is enabled. The former implements graph-mode execution (auto-defining placeholders dynamically), the latter eager execution.</p>
<p>The main difference between <code class="highlighter-rouge">DynamicTFPolicy</code> and <code class="highlighter-rouge">EagerTFPolicy</code> is how many times they call the functions passed in. In either case, a <code class="highlighter-rouge">model_fn</code> is invoked once to create a Model class. However, functions that involve tensor operations are either called once in graph mode to build a symbolic computation graph, or multiple times in eager mode on actual tensors. In the following figures we show how these operations work together in blue and orange:</p>
<p style="text-align:center;"> <!-- <img decoding="async" src="https://cdn-images-1.medium.com/max/800/1*4RKNH6Zt4-P82ZeoETKVog.png"> --> <img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img10.png" />  </p>
<p style="text-align:center;"> <!-- <img decoding="async" src="https://cdn-images-1.medium.com/max/800/0*1vTJnByqwRvPoyBV"> --> <img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img11.png" /> <br /> <i> Overview of a generated EagerTFPolicy. The policy passes the environment state through model.forward(), which emits output logits. The model output parameterizes a probability distribution over actions (“ActionDistribution”), which can be used when sampling actions or training. The loss function operates over batches of experiences. The model can provide additional methods such as a value function (light orange) or other methods for computing Q values, etc. (not shown) as needed by the loss function. </i> </p>
<p>This policy object is all RLlib needs to launch and scale RL training. Intuitively, this is because it encapsulates how to compute actions and improve the policy. External state such as that of the environment and RNN hidden state is managed externally by RLlib, and does not need to be part of the policy definition. The policy object is used in one of two ways depending on whether we are computing rollouts or trying to improve the policy given a batch of rollout data:</p>
<p style="text-align:center;"> <!-- <img decoding="async" src="https://cdn-images-1.medium.com/max/800/0*-35jGBA7Gha9WnOA"> --> <img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img12.gif" /> <br /> <i> <b>Inference:</b> Forward pass to compute a single action. This only involves querying the model, generating an action distribution, and sampling an action from that distribution. In eager mode, this involves calling action_fn <a href="https://github.com/ray-project/ray/blob/03a1b758526b2699a21e44a932bb2abdfe636f2b/rllib/agents/dqn/simple_q_policy.py#L111" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"> DQN example of an action sampler</a>, which creates an action distribution / action sampler as relevant that is then sampled from. </i> </p>
<p style="text-align:center;"> <!-- <img decoding="async" src="https://cdn-images-1.medium.com/max/800/0*w4cJ0KPTM8QPX5Ex"> --> <img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img13.gif" /> <br /> <i> <b>Training:</b> Forward and backward pass to learn on a batch of experiences. In this mode, we call the loss function to generate a scalar output which can be used to optimize the model variables via SGD. In eager mode, both action_fn and loss_fn are called to generate the action distribution and policy loss respectively. Note that here we don’t show differentiation through action_fn, but this does happen in algorithms such as DQN. </i> </p>
<h3 id="loose-ends-state-management">Loose Ends: State Management</h3>
<p>RL training inherently involves a lot of state. If algorithms are defined using pure functions, where is the state held? In most cases it can be managed automatically by the framework. There are three types of state that need to be managed in RLlib:</p>
<ol>
<li><strong>Environment state</strong>: this includes the current state of the environment and any recurrent state passed between policy steps. RLlib manages this internally in its <a href="https://github.com/ray-project/ray/blob/master/rllib/evaluation/rollout_worker.py" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">rollout worker</a> implementation.</li>
<li><strong>Model state</strong>: these are the policy parameters we are trying to learn via an RL loss. These variables must be accessible and optimized in the same way for both graph and eager mode. Fortunately, <a href="https://www.tensorflow.org/guide/keras" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Keras</a> models can be used in either mode. RLlib provides a <a href="https://ray.readthedocs.io/en/latest/rllib-models.html#tensorflow-models" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">customizable model class (TFModelV2)</a> based on the object-oriented Keras style to hold policy parameters.</li>
<li><strong>Training workflow state</strong>: state for managing training, e.g., the annealing schedule for various hyperparameters, steps since last update, and so on. RLlib lets algorithm authors add <a href="https://github.com/ray-project/ray/blob/b520f6141ecdd54496b0c26106f3df4442a5f91e/rllib/agents/ppo/ppo_policy.py#L284" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">mixin classes</a> to policies that can hold any such extra variables.</li>
</ol>
<h3 id="loose-ends-eageroverhead">Loose ends: Eager Overhead</h3>
<p>Next we investigate RLlib’s eager mode performance with <a href="https://www.tensorflow.org/beta/tutorials/eager/tf_function" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">eager tracing</a> on or off. As shown in the below figure, tracing greatly improves performance. However, the tradeoff is that Python operations such as print may not be called each time. For this reason, tracing is off by default in RLlib, but can be enabled with “eager_tracing”: True. In addition, you can also set “no_eager_on_workers” to enable eager only for learning but disable it for inference:</p>
<p style="text-align:center;"> <!-- <img decoding="async" src="https://cdn-images-1.medium.com/max/800/0*YvplJKQocQXjclg5"> --> <img decoding="async" src="https://bair.berkeley.edu/static/blog/functional/img14.png" />  </p>
<p>Eager inference and gradient overheads measured using <code class="highlighter-rouge">rllib train --run=PG --env=&lt;env&gt; [ --eager [ --trace]]</code> on a laptop processor. With tracing off, eager imposes a significant overhead for small batch operations. However it is often as fast or faster than graph mode when tracing is enabled.</p>
<h3 id="conclusion">Conclusion</h3>
<p>To recap, in this blog post we propose using ideas from functional programming to simplify the development of RL algorithms. We implement and validate these ideas in RLlib. Beyond making it easy to support new features such as eager execution, we also find the functional paradigm leads to substantially more concise and understandable code. Try it out yourself with <code class="highlighter-rouge">pip install ray[rllib]</code> or by checking out the <a href="https://ray.readthedocs.io/en/latest/rllib.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">docs</a> and <a href="https://github.com/ray-project/ray/tree/master/rllib" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">source code</a>.</p>
<p>If you’re interested in helping improve RLlib, we’re also <a href="https://jobs.lever.co/anyscale" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">hiring</a>.</p>
<p>This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Deep dynamics models for dexterous manipulation</title>
		<link>https://robohub.org/deep-dynamics-models-for-dexterous-manipulation/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Thu, 03 Oct 2019 21:45:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/deep-dynamics-models-for-dexterous-manipulation/</guid>

					<description><![CDATA[
Figure 1: Our approach (PDDM) can efficiently and effectively learn complex
dexterous manipulation skills in both simulation and the real world. Here, the
learned model is able to control the 24-DoF Shadow Hand to rotate two
free-floating Baoding...]]></description>
										<content:encoded><![CDATA[<img decoding="async" src="https://robohub.org/wp-content/uploads/2019/09/DeepLearningManipulation.png" alt="" width="900" height="321" class="aligncenter size-full wp-image-147915" srcset="https://robohub.org/wp-content/uploads/2019/09/DeepLearningManipulation.png 900w, https://robohub.org/wp-content/uploads/2019/09/DeepLearningManipulation-425x152.png 425w, https://robohub.org/wp-content/uploads/2019/09/DeepLearningManipulation-768x274.png 768w" sizes="(max-width: 900px) 100vw, 900px" />
<p><strong>By Anusha Nagabandi  </strong>  </p>
<p>Dexterous manipulation with multi-fingered hands is a grand challenge in robotics: the versatility of the human hand is as yet unrivaled by the capabilities of robotic systems, and bridging this gap will enable more general and capable robots. Although some real-world tasks (like picking up a television remote or a screwdriver) can be accomplished with simple parallel jaw grippers, there are countless tasks (like functionally using the remote to change the channel or using the screwdriver to screw in a nail) in which dexterity enabled by redundant degrees of freedom is critical. In fact, dexterous manipulation is <a href="http://www-cdr.stanford.edu/Touch/publications/okamura_icra00.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">defined</a> as being object-centric, with the goal of controlling object movement through precise control of forces and motions — something that is not possible without the ability to simultaneously impact the object from multiple directions. For example, using only two fingers to attempt common tasks such as opening the lid of a jar or hitting a nail with a hammer would quickly encounter the challenges of slippage, complex contact forces, and underactuation. Although dexterous multi-fingered hands can indeed enable flexibility and success of a wide range of manipulation skills, many of these more complex behaviors are also notoriously difficult to control: They require finely balancing contact forces, breaking and reestablishing contacts repeatedly, and maintaining control of unactuated objects. Success in such settings requires a sufficiently dexterous hand, as well as an intelligent policy that can endow such a hand with the appropriate control strategy. We study precisely this in our work on Deep Dynamics Models for Learning Dexterous Manipulation.</p>
<p>  <span id="more-147497"></span>  </p>
<p><!-- <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image16.gif" width="600"> -->  </p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image12.gif" width="250" /> <br /> <i> Figure 1: Our approach (PDDM) can efficiently and effectively learn complex dexterous manipulation skills in both simulation and the real world. Here, the learned model is able to control the 24-DoF Shadow Hand to rotate two free-floating Baoding balls in the palm, using just 4 hours of real-world data with no prior knowledge/assumptions of system or environment dynamics. </i> </p>
<p>Common approaches for control include modeling the system as well as the relevant objects in the environment, planning through this model to produce reference trajectories, and then developing a controller to actually achieve these plans. However, the success and scale of these approaches have been restricted thus far due to their need for accurate modeling of complex details, which is especially difficult for such contact-rich tasks that call for precise fine-motor skills. Learning has thus become a popular approach, offering a promising data-driven method for directly learning from collected data rather than requiring explicit or accurate modeling of the world. Model-free reinforcement learning (RL) methods, in particular, have been shown to learn policies that achieve <a href="https://arxiv.org/pdf/1801.01290.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">good</a> <a href="http://www.jmlr.org/papers/volume17/15-522/15-522.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">performance</a> on <a href="https://arxiv.org/pdf/1808.00177.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">complex</a> tasks; however, we will show that these state-of-the-art algorithms struggle when a high degree of flexibility is required, such as moving a pencil to follow <em>arbitrary</em> user-specified strokes, instead of a fixed one. Model-free methods also require large amounts of data, often making them infeasible for real-world applications. Model-based RL methods, on the other hand, can be much more efficient, but have not yet been scaled up to similarly complex tasks. Our work aims to push the boundary on this task complexity, enabling a dexterous manipulator to turn a valve, reorient a cube in-hand, write arbitrary motions with a pencil, and rotate two Baoding balls around the palm. We show that our method of online planning with deep dynamics models (PDDM) addresses both of the aforementioned limitations: Improvements in learned dynamics models, together with improvements in online model-predictive control, can indeed enable efficient and effective learning of flexible contact-rich dexterous manipulation skills — and that too, on a 24-DoF anthropomorphic hand in the real world, using ~4 hours of purely real-world data to coordinate multiple free-floating objects.</p>
<h1 id="method-overview">Method Overview</h1>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image10.png" width="" /> <br /> <i> Figure 2: Overview of our PDDM algorithm for online planning with deep dynamics models. </i> </p>
<p>Learning complex dexterous manipulation skills on a real-world robotic system requires an algorithm that is (1) data-efficient, (2) flexible, and (3) general-purpose. First, the method must be efficient enough to learn tasks in just a few hours of interaction, in contrast to <a href="https://openai.com/blog/learning-dexterity/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">methods</a> that <a href="http://openaccess.thecvf.com/content_CVPR_2019/papers/James_Sim-To-Real_via_Sim-To-Sim_Data-Efficient_Robotic_Grasping_via_Randomized-To-Canonical_Adaptation_Networks_CVPR_2019_paper.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">utilize</a> <a href="https://arxiv.org/abs/1610.04286" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">simulation</a> and require hundreds of hours, days, or even years to learn. Second, the method must be flexible enough to handle a variety of tasks, so that the same model can be used to perform various different tasks. Third, the method must be general and make relatively few assumptions: It should not require a known model of the system, which can be very difficult to obtain for arbitrary objects in the world.</p>
<p>To this end, we adopt a model-based reinforcement learning approach for dexterous manipulation. Model-based RL methods work by learning a predictive model of the world, which predicts the next state given the current state and action. Such algorithms are more efficient than model-free learners because every trial provides rich supervision: even if the robot does not succeed at performing the task, it can use the trial to learn more about the physics of the world. Furthermore, unlike model-free learning, model-based algorithms are “off-policy,” meaning that they can use any (even old) data for learning. Typically, it is believed that this efficiency of model-based RL algorithms comes at a price: since they must go through this intermediate step of learning the model, they might not perform as well at convergence as model-free methods, which more directly optimize the reward. However, our simulated comparative evaluations show that our model-based method actually performs better than model-free alternatives when the desired tasks are very diverse (e.g., writing different characters with a pencil). This separation of modeling from control allows the model to be easily reused for different tasks – something that is not as straightforward with learned policies.</p>
<p>Our complete method (Figure 2), consists of learning a predictive model of the environment (denoted $f_\theta(s,a) = s’$), which can then be used to control the robot by planning a course of action at every time step through a sampling-based planning algorithm. Learning proceeds as follows: data is iteratively collected by attempting the task using the latest model, updating the model using this experience, and repeating. Although the basic design of our model-based RL algorithms has been explored in prior work, the particular design decisions that we made were crucial to its performance. We utilize an ensemble of models, which accurately fits the dynamics of our robotic system, and we also utilize a more powerful sampling-based planner that preferentially samples temporally correlated action sequences as well as performs reward-weighted updates to the sampling distribution. Overall, we see effective learning, a nice separation of modeling and control, and an intuitive mechanism for iteratively learning more about the world while simultaneously reasoning at each time step about what to do.</p>
<h1 id="baoding-balls">Baoding Balls</h1>
<p>For a true test of dexterity, we look to the task of <a href="https://mindworks.org/blog/history-benefits-and-uses-of-meditation-balls/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Baoding balls</a>. Also referred to as Chinese relaxation balls, these two free-floating spheres must be rotated around each other in the palm. Requiring both dexterity and coordination, this task is commonly used for improving finger coordination, relaxing muscular tensions, and recovering muscle strength and motor skills after surgery. Baoding behaviors evolve in the high dimensional workspace of the hand and exhibit contact-rich (finger-finger, finger-ball, and ball-ball) interactions that are hard to reliably capture, either analytically or even in a physics simulator. Successful baoding behavior on physical hardware requires not only learning about these interactions via real world experiences, but also effective planning to find precise and coordinated maneuvers while avoiding task failure (e.g., dropping the balls).</p>
<p>For our experiments, we use the ShadowHand — a 24-DoF five-fingered anthropomorphic hand. In addition to ShadowHand’s inbuilt proprioceptive sensing at each joint, we use a 280&#215;180 RGB stereo image pair that is fed into a separately pretrained tracker to produce 3D position estimates for the two Baoding balls. To enable continuous experimentation in the real world, we developed an automated reset mechanism (Figure 3) that consists of a ramp and an additional robotic arm: The ramp funnels the dropped Baoding balls to a specific position and then triggers the 7-DoF Franka-Emika arm to use its parallel jaw gripper to pick them up and return them to the ShadowHand’s palm to resume training. We note that the entire training procedure is performed using the hardware setup described above, without the aid of any simulation data.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image9.gif" width="" /> <br /> <i> Figure 3: Automated reset procedure, where the Franka-Emika arm gathers and resets the Baoding Balls, in order for the ShadowHand to continue its training. </i> </p>
<p>During the initial phase of the learning, the hand continues to drop both balls, since that is the very likely outcome before it knows how to solve the task. Later, it learns to keep the balls in the palm to avoid the penalty incurred due to dropping. As learning improves, progress in terms of half-rotations start to emerge around 30 minutes of training. Getting the balls past this 90-degree orientation is a difficult maneuver, and PDDM spends a moderate amount of time here: To get past this point, notice the transition that must happen (in the 3rd video panel of Figure 4), from first controlling the objects with the pinky, and then controlling them indirectly through hand motion, and finally getting to control them with the thumb. By ~2 hours, the hand can reliably make 90-degree turns, frequently make 180-degree turns, and sometimes even make turns with multiple rotations.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image15.gif" height="280" width="200" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image8.gif" height="280" width="200" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image13.gif" height="280" width="200" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image12.gif" height="280" width="200" /> <br /> <i> Figure 4: Training progress on the ShadowHand hardware. From left to right: 0-0.25 hours, 0.25-0.5 hours, 0.5-1.5 hours, ~2 hours. </i> </p>
<h1 id="simulated-tasks">Simulated Tasks</h1>
<p>Although we presented the PDDM algorithm in light of the Baoding task, it is very generic, and we show it below in Figure 5 working on a suite of simulated dexterous manipulation tasks. These tasks illustrate various challenges presented by contact-rich dexterous manipulation tasks — high dimensionality of the hand, intermittent contact dynamics involving hand and objects, prevalence of constraints that must be respected and utilized to effectively manipulate objects, and catastrophic failures from dropping objects from the hand.  These tasks not only require precise understanding of the rich contact interactions but also require carefully coordinated and planned movements.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image2.gif" width="200" height="200" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image7.gif" width="200" height="200" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image4.gif" width="200" height="200" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image6.gif" width="200" height="200" /> <br /> <i> Figure 5: Result of PDDM solving simulated dexterous manipulation tasks. From left to right: 9 DOF D&#8217;Claw turning valve to random (green) targets (~20 min of data), 16 dof D&#8217;Hand pulling a weight via the manipulation of a flexible rope (~1 hour of data), 24 DOF ShadowHand performing in-hand reorientation of a free-floating cube to random (shown) targets (~1 hour of data), 24 DOF ShadowHand following desired trajectories with tip of a free-floating pencil (~1-2 hours of data). Note that the amount of data is measured in terms of the real-world equivalent (e.g., 100 data points where each step represents 0.1 seconds would represent 10 seconds worth of data). </i> </p>
<h2 id="model-reuse">Model Reuse</h2>
<p>Since PDDM learns dynamics models as opposed to task-specific policies or policy-conditioned value functions, a given model can then be reused when planning for different but related tasks. In Figure 6 below, we demonstrate that the model trained for the Baoding task of performing counterclockwise rotations (left) can be repurposed to move a single ball to a goal location in the hand (middle) or to perform clockwise rotations (right) instead of the learned counterclockwise ones.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image5.gif" height="230" width="270" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image1.gif" height="230" width="270" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image11.gif" height="230" width="270" /> <br /> <i> Figure 6: Model reuse on simulated tasks. Left: train model on CCW Baoding task. Middle: reuse that model for go-to single location task. Right: reuse that same model for CW Baoding task. </i> </p>
<h2 id="flexibility">Flexibility</h2>
<p>We study the flexibility of PDDM by experimenting with handwriting, where the base of the hand is fixed and arbitrary characters need to be written through the coordinated movement of the fingers and wrist. Although even writing a fixed trajectory is challenging, we see that writing arbitrary trajectories requires a degree of flexibility and coordination that is exceptionally challenging for prior methods. PDDM’s separation of modeling and task-specific control allows for generalization across behaviors, as opposed to discovering and memorizing the answer to a specific task/movement. In Figure 7 below, we show PDDM’s handwriting results that were trained on random paths for the green dot but then tested in a zero-shot fashion to write numerical digits.</p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image14.gif" width="360" /> <img decoding="async" src="https://bair.berkeley.edu/static/blog/deep-dynamics/image3.gif" width="360" /> <br /> <i> Figure 7: Flexibility of the learned handwriting model, which was trained to follow random paths of the green dot, but shown here to write some digits. </i> </p>
<h1 id="future-directions">Future Directions</h1>
<p>Our results show that PDDM can be used to learn challenging dexterous manipulation tasks, including controlling free-floating objects, agile finger gaits for repositioning objects in the hand, and precise control of a pencil to write user-specified strokes. In addition to testing PDDM on our simulated suite of tasks to analyze various algorithmic design decisions as well as to perform comparisons to other state-of-the-art model-based and model-free algorithms, we also show PDDM learning the Baoding Balls task on a real-world 24-DoF anthropomorphic hand using just a few hours of entirely real-world interaction. Since model-based techniques do indeed show promise on complex tasks, exciting directions for future work would be to study methods for planning at different levels of abstraction to enable success on sparse-reward or long-horizon tasks, as well as to study the effective integration of additional sensing modalities, such as vision and touch, into these models to better understand the world and expand the boundaries of what our robots can do. Can our robotic hand braid someone’s hair? Crack an egg and carefully handle the shell? Untie a knot? Button up all the buttons of a shirt? Tie shoelaces? With the development of models that can understand the world, along with planners that can effectively use those models, we hope the answer to all of these questions will become ‘yes.’</p>
<h2 id="acknowledgements">Acknowledgements</h2>
<p>This <a href="https://sites.google.com/view/pddm/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">work</a> was done at Google Brain, and the authors are Anusha Nagabandi, Kurt Konoglie, Sergey Levine, and Vikash Kumar.  The authors would also like to thank Michael Ahn for his frequent software and hardware assistance, and Sherry Moore for her work on setting up the drivers and code for working with our ShadowHand.</p>
<p> This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission. </p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Sample efficient evolutionary algorithm for analog circuit design</title>
		<link>https://robohub.org/sample-efficient-evolutionary-algorithm-for-analog-circuit-design/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Thu, 03 Oct 2019 19:00:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://bair.berkeley.edu/blog/2019/09/26/circuits/</guid>

					<description><![CDATA[<p>In this post, we share some recent promising results regarding the applications
of Deep Learning in analog IC design. While this work targets a specific
application, the proposed methods can be used in other black box optimization
problems where the environment lacks a cheap/fast evaluation procedure.</p>

&#60;!--
<img src="https://bair.berkeley.edu/static/blog/circuits/1.png" width="600">
--&#62;

<p>
<img src="https://bair.berkeley.edu/static/blog/circuits/title_image_v02.svg" width="600"><br></p>

<p>So let&#8217;s break down how the analog IC design process is usually done, and then
how we incorporated deep learning to ease the flow.</p>

<!--more-->

<p>The intent of analog IC design is to build a physical manufacturable circuit
that processes electrical signals in the analog domain, despite all sorts of
noise sources that may affect the fidelity of signals. Usually analog circuit
design starts off with topology selection. Generically speaking, engineers
usually come up with topology of certain blocks and try to size them such that
after putting them together the entire system behaves in a certain way and
satisfies some figures of merit. There are certain levels of simulations and
tests that need to take place to verify that the system will work before
manufacturing. At the lowest level engineers do their design using their
intuition and equations and then simulate and make the corresponding changes
until they converge to a working design. The more accurate the simulation, the
more time it takes to run. Unfortunately, in recent advanced technologies the
large disparity between post-layout simulation (physical realization of the
circuit) and schematic simulation (circuit concept) requires designers to be
aware of the parasitic effects due to the way the circuit is physically
implemented. This basically means that simulations take longer, and on top of
that more manual iterations are needed.</p>

<p>Now that we understand the environment let&#8217;s give a brief overview of past
attempts to automate some parts of this process.</p>

<p>Historically people have tried different degrees of automation in different
parts of the design flow (a complete ancient survey can be found <a href="https://www.wiley.com/en-us/Computer+Aided+Design+of+Analog+Integrated+Circuits+and+Systems-p-9780471227823" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">in this
textbook</a>, and more recent work includes <a href="https://ieeexplore.ieee.org/document/8116661" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Bayesian optimization</a> and
<a href="https://arxiv.org/abs/1812.02734" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RL</a>). We are mostly interested in approaches that find the optimal sizing for a
given topology to satisfy a collection of metrics (a constraint satisfaction
problem).</p>

<p>Some people have derived analytical formulations for behavior of circuits and turned the problem into optimization over these analytical expressions (for example by expressing gain as an analytical function of transistor sizes). However, as was mentioned earlier, today, even schematic simulations can differ from their layout counterpart. So those ancient approaches lost their attractiveness very early on due to this important
drawback.</p>

<p>Some other approaches tried to use simulations and modeled the problem as a
black box optimization. A lot of them showed success in simple circuits in
schematic-based simulations, but again could not scale very well to larger
circuits and layout exploration.  The bottom line till now is that a lot of
iterations are needed and long simulations make it even more difficult and
cumbersome.</p>

<p>On a side note, population based black-box optimization algorithms achieve a pretty good performance, in terms of the final output&#8217;s quality.
For example, <a href="https://arxiv.org/abs/1703.03864" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this paper</a> shows a setting where RL agents are trained in a parallelized fashion using scalable evolutionary algorithms.  The problem is that they are insanely sample inefficient (despite
being parallelizable) and their exploration strategy is mostly stochastic with
no &#8220;real&#8221; guidance. The problem with circuit design is that tools available for
simulation are not highly parallelizable or very expensive to parallelize. In this work we have proposed a new way to make them more sample efficient.</p>

<!--
Let&#8217;s formalize the problem a bit. Let&#8217;s say we want to design an amplifier
with a given topology for a given gain $A_0$ and bandwidth $W_0$. We want to find
&#8220;optimum&#8221; sizes for components of the circuit such that the gain and bandwidth
are larger than let&#8217;s say $A_0$ and $W_0$, respectively.
-->

<p>Let&#8217;s formalize the problem a bit. Let&#8217;s say we want to design an amplifier
with a given topology for a gain larger than $A_0$ and a bandwidth larger than
$BW_0$. We want to find the &#8220;optimum&#8221; sizes for the components of the circuit such that they satisfy these performance constraints.
We can formulate the problem as minimizing a scalar cost function equal to the sum of relative errors to the required specification. More concretely, in this example:</p>

<p>And we get the $A(x)$ and $BW(x)$ values after simulation.</p>

<p>To state it in a more general form:</p>

<p>Where $x$ presents the geometric parameters in the circuit topology and</p>

<p>represents the normalized spec error for designs that do not satisfy constraint
, or zero if they do. $c_i$ denotes the value of constraint $i$ at
input $x$, and is evaluated using a simulation framework.  denotes the
optimal value.  Intuitively this cost function is only accounting for the
normalized error from the unsatisfied constraints, and $w_i$ is the tuning
factor, determined by the designer, which controls prioritizing one metric over
another if the design is infeasible.</p>

<p>In principle, we can start optimizing $cost(x)$ by using evolutionary
algorithms (a great intro found <a href="https://medium.com/sigmoid/https-medium-com-rishabh-anand-on-the-origin-of-genetic-algorithms-fc927d2e11e0" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>). In fact we used <a href="https://deap.readthedocs.io/en/master/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">deap</a> to
implement a baseline version of genetic algorithm to our problem. The issue is that
there is no &#8220;real&#8221; intelligence in exploration, and therefore, it takes a lot of
expensive simulations to find a solution, even for the simplest design
problems. To put it in perspective let&#8217;s work with a 3D contrived example,
let&#8217;s say the current population includes $x_1$, $x_2$, and $x_3$ among which
none satisfy all specs under consideration.</p>

&#60;!--
<p style="text-align:center">
    <img src="https://bair.berkeley.edu/static/blog/circuits/2.png" width="300">
    <br>
</p>
--&#62;

<p>Now let&#8217;s say our chance hits and we produce the following $y$ samples from the
old population.</p>

&#60;!--
<p style="text-align:center">
    <img src="https://bair.berkeley.edu/static/blog/circuits/3.png" width="300">
    <br>
</p>

<p style="text-align:center">
    <img src="https://bair.berkeley.edu/static/blog/circuits/4.png" width="300">
    <br>
</p>

<p style="text-align:center">
    <img src="https://bair.berkeley.edu/static/blog/circuits/5.png" width="250">
    <br>
</p>
--&#62;

<p>We know the performance of $x$ samples and to know performance of $y$ samples
we need to run simulations (this is the part which can potentially be extremely
slow).</p>

<p>Now let&#8217;s say after simulation we sort samples by their performance and get the
following next generation of population.</p>

&#60;!--
<p style="text-align:center">
    <img src="https://bair.berkeley.edu/static/blog/circuits/6.png" width="200">
    <br>
</p>
--&#62;

<p>Observe that after this &#8220;unlucky&#8221; iteration $y_2$ and $y_3$ got eliminated. In
our proposed method, we devised a model to predict the performance before
simulation and only simulate those samples which have better predicted
performance. So in our contrived example we&#8217;ll predict whether new designs have
better performance than $x_1$ and if so, we simulate them. If our prediction is
accurate (or almost accurate) we waste fewer simulations. On the other hand,
if we do not make accurate decisions, we could either approve samples that
do not show high quality after simulation or reject designs that
should have been accepted.</p>

<p>Next we&#8217;ll describe a model that was able to achieve acceptable results
utilizing this idea. All implementations details can be found in <a href="https://github.com/kouroshHakha/bag_deep_ckt" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">the GitHub
code</a>.</p>

<h1>Architecture Choice for the Model</h1>

<p>This model has to have two distinct characteristics:</p>

<ol><li>
    <p>It has to be able to express how good/bad a new, not simulated sample is
compared to individuals in the current population.</p>
  </li>
  <li>
    <p>It should be able to generalize well, given very limited number of training
samples (potentially 100-300 accurate simulations). However, it should not be
biased towards a specific region in space and get stuck in a local optimum.</p>
  </li>
</ol><p>The first potential candidate is a regression model which predicts the cost
value and uses that prediction to determine whether to simulate a design.  The
cost function that the network tries to approximate can be a non-convex and
ill-conditioned one. Thus, from a limited number of samples it is very unlikely
that it would generalize well to unseen data. Moreover, the cost function
captures too much information from a single scalar number, so it would be hard
to train, given a small number of training points and then expect it to
generalize well.</p>

<p>Another option is to predict the value of each metric (i.e. gain, bandwidth,
etc.). While the individual metric behaviour can be smoother than the cost
function, predicting the actual metric value is unnecessary, since we are
simply attempting to predict whether a new design is superior to some other
design. Therefore, instead of predicting metric values exactly, the model can
take two designs and predict only which design performs better in each
individual metric.</p>

<p>We know that there are certain patterns in circuit parameters that make some
metrics better than others (for example upsizing all transistors make the
amplifier faster but costs more power). We use those parameters as the input of
our network. Moreover, by forming a model that takes in pairs of designs
instead of a single design we effectively expand our training sample size (a
population of size 100 has ~5000 pairs). Despite the fact that neural net&#8217;s
input space is larger, in practice, training is easier.</p>

<p>The bottom figure shows the structure of the neural net modeling the oracle
simulator used in this approach. This model takes in two sets
of parameters (one for Design A and one for B); extracts some useful features
$(f_1, \ldots, f_k)$; rearranges the order of features (the reason for this
will become clear shortly); and feeds the cascaded/rearranged feature vectors
to independent neural networks to predict the superiority of Design A to Design
B for a given spec. The dedicated superiority-predicting neural nets share the
same architecture across all specs, but are parameterized separately.  The
output is interpreted as the probability that Design A is better than Design B
in $\text{spec}_1$ or $\text{spec}_2$ and so on.</p>

<p>So during inference, the parameters of the query design are going to be fed to
Design B, and the reference design to Design A. Then we predict how the query
design will do compared to the reference design. Then we use some heuristic to
decide whether to simulate or ignore the new design.</p>

<p>
    <img src="https://bair.berkeley.edu/static/blog/circuits/7.png" width="600"><br></p>

<p>One other subtle constraint on the model is that there should be no
contradiction in the predicted probabilities depending on the order by which
the inputs were fed in. For example if we feed in A and B, and A is predicted
to have a better gain than B with probability 0.8, if we swap the order of A
and B the probability should become 0.2 by construction of the network.
This requires a particular symmetry in the network architecture. For
more info on this refer to the <a href="https://arxiv.org/abs/1907.10515" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a>, but here is a code snippet that
shows how we make sure a layer behaves like that by construction.</p>

<div><div><pre><code><span>weight_elements</span> <span>=</span> <span>tf</span><span>.</span><span>get_variable</span><span>(</span><span>name</span><span>=</span><span>'W'</span><span>,</span> <span>shape</span><span>=</span><span>[</span><span>input_data</span><span>.</span><span>shape</span><span>[</span><span>1</span><span>]</span><span>//</span><span>2</span><span>,</span> <span>layer_dim</span><span>],</span> <span>initializer</span><span>=</span><span>tf</span><span>.</span><span>random_normal_initializer</span><span>)</span>
<span>bias_elements</span> <span>=</span> <span>tf</span><span>.</span><span>get_variable</span><span>(</span><span>name</span><span>=</span><span>'b'</span><span>,</span> <span>shape</span><span>=</span><span>[</span><span>layer_dim</span><span>//</span><span>2</span><span>],</span> <span>initializer</span><span>=</span><span>tf</span><span>.</span><span>zeros_initializer</span><span>)</span>
<span>Weight</span> <span>=</span> <span>tf</span><span>.</span><span>concat</span><span>([</span><span>weight_elements</span><span>,</span> <span>weight_elements</span><span>[::</span><span>-</span><span>1</span><span>,</span> <span>::</span><span>-</span><span>1</span><span>]],</span> <span>axis</span><span>=</span><span>0</span><span>,</span> <span>name</span><span>=</span><span>'Weights'</span><span>)</span>
<span>Bias</span> <span>=</span> <span>tf</span><span>.</span><span>concat</span><span>([</span><span>bias_elements</span><span>,</span> <span>bias_elements</span><span>[::</span><span>-</span><span>1</span><span>]],</span> <span>axis</span><span>=</span><span>0</span><span>,</span> <span>name</span><span>=</span><span>'Bias'</span><span>)</span>
</code></pre></div></div>

<p>The reason that the rearrangement happens is exactly this and the code below shows
how it&#8217;s actually done.</p>

<div><div><pre><code><span>features1</span> <span>=</span> <span>self</span><span>.</span><span>_feature_extraction_model</span><span>(</span><span>input1_norm</span><span>,</span> <span>name</span><span>=</span><span>'feat_model'</span><span>,</span> <span>reuse</span><span>=</span><span>False</span><span>)</span>
<span>features2</span> <span>=</span> <span>self</span><span>.</span><span>_feature_extraction_model</span><span>(</span><span>input2_norm</span><span>,</span> <span>name</span><span>=</span><span>'feat_model'</span><span>,</span> <span>reuse</span><span>=</span><span>True</span><span>)</span>
<span>input_features</span> <span>=</span> <span>tf</span><span>.</span><span>concat</span><span>([</span><span>features1</span><span>,</span> <span>features2</span><span>[:,</span> <span>::</span><span>-</span><span>1</span><span>]],</span> <span>axis</span><span>=</span><span>1</span><span>)</span>
</code></pre></div></div>

<p>To train the network, we construct all Design A and Design B permutations from
the buffer of previously simulated designs and label their comparison in each
metric. We then update network parameters with Adam optimizer, using sum of
cross-entropy loss for all metrics.</p>

<p>There is also another architecture choice that helped in the overall
convergence and that is using dropout layers to estimate the uncertainty in
predictions. During each query,  we randomly turn off 20% of the activations in
each layer 5 times, and average the output probabilities to remove the
uncertainty.</p>

<p>
    <img src="https://bair.berkeley.edu/static/blog/circuits/8.png" width="600"><br></p>

<p>Let&#8217;s describe the figure above. Our algorithm uses some underlying
evolutionary algorithm that has a broad exploration strategy. We used deap to
implement our algorithm&#8217;s evolutionary strategies (code found <a href="https://github.com/kouroshHakha/bag_deep_ckt/tree/master/deepckt/ea" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>).</p>

<p>Our algorithm starts by randomly sampling the design space (for let&#8217;s say 100
designs) and simulates all of them (this part takes some time). Then the
evolutionary algorithm proposes some offspring for future generations. We then
use the predictive model to &#8220;guess&#8221; whether the new proposed offspring is
better than some average design in the current population. So that when we
really simulate, it doesn&#8217;t get eliminated.</p>

<p>We should keep re-training the model as more samples are added to our database;
otherwise, the model gets biased to a specific region and we lose accuracy as
we progress. <a href="https://arxiv.org/abs/1011.0686" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DAgger</a> does this for imitation learning, however the big
difference here is that we can&#8217;t really re-label (simulate) rejected samples as
it defeats the purpose of short run time, so we only simulate those samples
which get approved by the model (either by mistake or correctly).</p>

<p>What&#8217;s that decision box at the output of the discriminator? That&#8217;s basically
telling the prediction process, the criteria to be used for discrimination.
Remember that output of model is whether a design is better than some other
design in individual metrics (not overall). The decision box in summary, keeps
track of the metrics we improved so far and the collection of metrics that
affect the cost objective the most. For more info on details please refer to
the paper, and the code.</p>

<h1>Finally, Experiments!</h1>

<p>We tested this methodology on a couple of circuits with different settings,
gradually making the complexity more similar to today&#8217;s analog circuit design
problems. The simplest relevant scenario is design of an opamp, verified in
schematic (with no layout) so that simulation is cheap and we can run an oracle
discriminator (based on real simulations rather than prediction of the model).</p>

<p>We compared that to our approach and the normal evolutionary algorithm.  The
details of the circuit are in the paper, but let&#8217;s look at the performance
curve to prove a point here.</p>

<p>
    <img src="https://bair.berkeley.edu/static/blog/circuits/9.png" width="500"><br></p>

<p>In this plot we are looking at the average cost function of the top 20
individuals in the current population as a function of number of iterations. In
each iteration we simulate 5 designs and evict the worst 5 designs (to keep the
evolving population size constant). If we use just the underlying evolutionary
algorithm (without any discriminations) we get the blue curve (with over 5000
simulations); however, our method with neural network discrimination produces
the green curve (with only 240 simulations). The orange curve shows the same
method if we use the simulator the do the discriminations. Conducting the
orange experiment requires simulation of all proposed samples and keeping top 5
that are actually better than design rank 20 in the current population, which
means a lot of simulations (3400 simulations). The gap between the orange and
green illustrates how loss of accuracy due to utilization of function
approximators affects the overall optimization performance. Reaching zero means
we have at least 20 designs satisfying all the specs.</p>

<p>We also experimented with an optical photonic receiver design - a circuit of
bigger size and longer simulations - to see if this method can be applied to
long post-layout simulations of designs which are more complex. In this
particular example we also showed that we can even design circuits with very high
level specifications. For this example the search space was of size $2.8 \times
10^{30}$ and to find the optimal designs we queried the discriminator 77487
times from which we only ran 435 simulations. This means that if we had to
simulate all of them we had to wait 300x longer.</p>

<h1>Future Directions</h1>

<p>Evaluating the performance of the algorithm in a quantitative way is still a
challenging problem. For the simple example above, we used comparisons
against an oracle. However, as our circuits get bigger and more complex,
running the oracle as a baseline quickly becomes unfeasible. How much certain
choices in the architecture affected the total performance is an open question.
We are looking into various figures of merit to account for both diversity and
quality of approved samples and compare models with different choices using
that.</p>

<p>In the results presented, we used full-accuracy post-layout simulations to
optimize the system. These alterations to the circuit sizing reflected both
fundamental circuit tradeoffs, as well as effects due to layout parasitics.
Because schematic simulation is significantly faster than post-layout
simulation, one direction for this work is to attempt to learn the circuit
tradeoffs from the schematic level simulation, and then refine the models based
on post-layout data.</p>

<hr><p>We refer the reader to the following paper for details:</p>

<ul><li><b><a href="https://arxiv.org/abs/1907.10515" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BagNet: Berkeley Analog Generator with Layout Optimizer Boosted with Deep Neural Networks</a></b></li>
</ul>]]></description>
										<content:encoded><![CDATA[<p><strong>By Kourosh Hakhamaneshi</strong></p>
<p>In this post, we share some recent promising results regarding the applications of Deep Learning in analog IC design. While this work targets a specific application, the proposed methods can be used in other black box optimization problems where the environment lacks a cheap/fast evaluation procedure.</p>
<p><span id="more-147002"></span>    <!-- <img decoding="async" src="https://bair.berkeley.edu/static/blog/circuits/1.png" width="600"> -->  </p>
<p style="text-align:center;"> <img decoding="async" src="https://bair.berkeley.edu/static/blog/circuits/title_image_v02.svg" width="600" />  </p>
<p>So let’s break down how the analog IC design process is usually done, and then how we incorporated deep learning to ease the flow.</p>
<p>  <!--more-->  </p>
<p>The intent of analog IC design is to build a physical manufacturable circuit that processes electrical signals in the analog domain, despite all sorts of noise sources that may affect the fidelity of signals. Usually analog circuit design starts off with topology selection. Generically speaking, engineers usually come up with topology of certain blocks and try to size them such that after putting them together the entire system behaves in a certain way and satisfies some figures of merit. There are certain levels of simulations and tests that need to take place to verify that the system will work before manufacturing. At the lowest level engineers do their design using their intuition and equations and then simulate and make the corresponding changes until they converge to a working design. The more accurate the simulation, the more time it takes to run. Unfortunately, in recent advanced technologies the large disparity between post-layout simulation (physical realization of the circuit) and schematic simulation (circuit concept) requires designers to be aware of the parasitic effects due to the way the circuit is physically implemented. This basically means that simulations take longer, and on top of that more manual iterations are needed.</p>
<p>Now that we understand the environment let’s give a brief overview of past attempts to automate some parts of this process.</p>
<p>Historically people have tried different degrees of automation in different parts of the design flow (a complete ancient survey can be found <a href="https://www.wiley.com/en-us/Computer+Aided+Design+of+Analog+Integrated+Circuits+and+Systems-p-9780471227823" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">in this textbook</a>, and more recent work includes <a href="https://ieeexplore.ieee.org/document/8116661" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Bayesian optimization</a> and <a href="https://arxiv.org/abs/1812.02734" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RL</a>). We are mostly interested in approaches that find the optimal sizing for a given topology to satisfy a collection of metrics (a constraint satisfaction problem).</p>
<p>Some people have derived analytical formulations for behavior of circuits and turned the problem into optimization over these analytical expressions (for example by expressing gain as an analytical function of transistor sizes). However, as was mentioned earlier, today, even schematic simulations can differ from their layout counterpart. So those ancient approaches lost their attractiveness very early on due to this important drawback.</p>
<p>Some other approaches tried to use simulations and modeled the problem as a black box optimization. A lot of them showed success in simple circuits in schematic-based simulations, but again could not scale very well to larger circuits and layout exploration.  The bottom line till now is that a lot of iterations are needed and long simulations make it even more difficult and cumbersome.</p>
<p>On a side note, population based black-box optimization algorithms achieve a pretty good performance, in terms of the final output’s quality. For example, <a href="https://arxiv.org/abs/1703.03864" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this paper</a> shows a setting where RL agents are trained in a parallelized fashion using scalable evolutionary algorithms.  The problem is that they are insanely sample inefficient (despite being parallelizable) and their exploration strategy is mostly stochastic with no “real” guidance. The problem with circuit design is that tools available for simulation are not highly parallelizable or very expensive to parallelize. In this work we have proposed a new way to make them more sample efficient.</p>
<p>  <!-- Let’s formalize the problem a bit. Let’s say we want to design an amplifier with a given topology for a given gain $A_0$ and bandwidth $W_0$. We want to find “optimum” sizes for components of the circuit such that the gain and bandwidth are larger than let’s say $A_0$ and $W_0$, respectively. -->  </p>
<p>Let’s formalize the problem a bit. Let’s say we want to design an amplifier with a given topology for a gain larger than $A_0$ and a bandwidth larger than $BW_0$. We want to find the “optimum” sizes for the components of the circuit such that they satisfy these performance constraints. We can formulate the problem as minimizing a scalar cost function equal to the sum of relative errors to the required specification. More concretely, in this example:</p>
<p>  <script type="math/tex; mode=display">% <![CDATA[ cost(x) = \frac{|A(x) - A_0|}{(A(x) + A_0)} \mathbf{1}(A(x) < A_0) + \frac{|BW(x) - BW_0|}{(BW(x) + BW_0)} \mathbf{1}({BW(x) < BW_0}) %]]&gt;</script>  </p>
<p>And we get the $A(x)$ and $BW(x)$ values after simulation.</p>
<p>To state it in a more general form:</p>
<p>  <script type="math/tex; mode=display">cost(x) = \sum_{i}{w_ip_i(x)}</script>  </p>
<p>Where $x$ presents the geometric parameters in the circuit topology and</p>
<p>  <script type="math/tex; mode=display">p_i(x)= \frac{|c_i - c_i^*|}{c_i+c_i^*}</script>  </p>
<p>represents the normalized spec error for designs that do not satisfy constraint <script type="math/tex">c_i^*</script>, or zero if they do. $c_i$ denotes the value of constraint $i$ at input $x$, and is evaluated using a simulation framework. <script type="math/tex">c_i^*</script> denotes the optimal value.  Intuitively this cost function is only accounting for the normalized error from the unsatisfied constraints, and $w_i$ is the tuning factor, determined by the designer, which controls prioritizing one metric over another if the design is infeasible.</p>
<p>In principle, we can start optimizing $cost(x)$ by using evolutionary algorithms (a great intro found <a href="https://medium.com/sigmoid/https-medium-com-rishabh-anand-on-the-origin-of-genetic-algorithms-fc927d2e11e0" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>). In fact we used <a href="https://deap.readthedocs.io/en/master/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">deap</a> to implement a baseline version of genetic algorithm to our problem. The issue is that there is no “real” intelligence in exploration, and therefore, it takes a lot of expensive simulations to find a solution, even for the simplest design problems. To put it in perspective let’s work with a 3D contrived example, let’s say the current population includes $x_1$, $x_2$, and $x_3$ among which none satisfy all specs under consideration.</p>
<p>  <!-- 

<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/circuits/2.png"     width="300">      </p>

 -->  <script type="math/tex; mode=display">\begin{align*}     x_1 = \begin{bmatrix}     2 \\     33 \\     43     \end{bmatrix},\     x_2 = \begin{bmatrix}     20 \\     3 \\     15     \end{bmatrix},\     x_3 =     \begin{bmatrix}     22 \\     15 \\     34     \end{bmatrix},\ \end{align*}</script>  </p>
<p>Now let’s say our chance hits and we produce the following $y$ samples from the old population.</p>
<p>  <!-- 

<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/circuits/3.png"     width="300">      </p>

  

<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/circuits/4.png"     width="300">      </p>

  

<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/circuits/5.png"     width="250">      </p>

 -->  <script type="math/tex; mode=display">\begin{align*}     \text{cross-over}: x_1, x_2 \to y_1 = \begin{bmatrix}     2 \\     3 \\     15     \end{bmatrix} \end{align*}</script>  <script type="math/tex; mode=display">\begin{align*}     \text{combination}: y_2 = 0.5x_1 + 0.5x_2 =  \begin{bmatrix}     11 \\     18 \\     29     \end{bmatrix} \end{align*}</script>  <script type="math/tex; mode=display">\begin{align*}     \text{Mutation}: y_3 =  \text{Mutate}(x_3)=  \begin{bmatrix}     50 \\     15 \\     10     \end{bmatrix} \end{align*}</script>  </p>
<p>We know the performance of $x$ samples and to know performance of $y$ samples we need to run simulations (this is the part which can potentially be extremely slow).</p>
<p>Now let’s say after simulation we sort samples by their performance and get the following next generation of population.</p>
<p>  <!-- 

<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/circuits/6.png"     width="200">      </p>

 -->  <script type="math/tex; mode=display">\begin{align*}     \begin{bmatrix}     x_1 \\     y_1 \\     x_2 \\     x_3 \\     y_2 \\     y_3 \\     \end{bmatrix} \to     \begin{bmatrix}     x_1 \\     y_1 \\     x_2 \\     \end{bmatrix} \end{align*}</script>  </p>
<p>Observe that after this “unlucky” iteration $y_2$ and $y_3$ got eliminated. In our proposed method, we devised a model to predict the performance before simulation and only simulate those samples which have better predicted performance. So in our contrived example we’ll predict whether new designs have better performance than $x_1$ and if so, we simulate them. If our prediction is accurate (or almost accurate) we waste fewer simulations. On the other hand, if we do not make accurate decisions, we could either approve samples that do not show high quality after simulation or reject designs that should have been accepted.</p>
<p>Next we’ll describe a model that was able to achieve acceptable results utilizing this idea. All implementations details can be found in <a href="https://github.com/kouroshHakha/bag_deep_ckt" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">the GitHub code</a>.</p>
<h1 id="architecture-choice-for-the-model">Architecture Choice for the Model</h1>
<p>This model has to have two distinct characteristics:</p>
<ol>
<li>
<p>It has to be able to express how good/bad a new, not simulated sample is compared to individuals in the current population.</p>
</li>
<li>
<p>It should be able to generalize well, given very limited number of training samples (potentially 100-300 accurate simulations). However, it should not be biased towards a specific region in space and get stuck in a local optimum.</p>
</li>
</ol>
<p>The first potential candidate is a regression model which predicts the cost value and uses that prediction to determine whether to simulate a design.  The cost function that the network tries to approximate can be a non-convex and ill-conditioned one. Thus, from a limited number of samples it is very unlikely that it would generalize well to unseen data. Moreover, the cost function captures too much information from a single scalar number, so it would be hard to train, given a small number of training points and then expect it to generalize well.</p>
<p>Another option is to predict the value of each metric (i.e. gain, bandwidth, etc.). While the individual metric behaviour can be smoother than the cost function, predicting the actual metric value is unnecessary, since we are simply attempting to predict whether a new design is superior to some other design. Therefore, instead of predicting metric values exactly, the model can take two designs and predict only which design performs better in each individual metric.</p>
<p>We know that there are certain patterns in circuit parameters that make some metrics better than others (for example upsizing all transistors make the amplifier faster but costs more power). We use those parameters as the input of our network. Moreover, by forming a model that takes in pairs of designs instead of a single design we effectively expand our training sample size (a population of size 100 has ~5000 pairs). Despite the fact that neural net’s input space is larger, in practice, training is easier.</p>
<p>The bottom figure shows the structure of the neural net modeling the oracle simulator used in this approach. This model takes in two sets of parameters (one for Design A and one for B); extracts some useful features $(f_1, \ldots, f_k)$; rearranges the order of features (the reason for this will become clear shortly); and feeds the cascaded/rearranged feature vectors to independent neural networks to predict the superiority of Design A to Design B for a given spec. The dedicated superiority-predicting neural nets share the same architecture across all specs, but are parameterized separately.  The output is interpreted as the probability that Design A is better than Design B in $\text{spec}_1$ or $\text{spec}_2$ and so on.</p>
<p>So during inference, the parameters of the query design are going to be fed to Design B, and the reference design to Design A. Then we predict how the query design will do compared to the reference design. Then we use some heuristic to decide whether to simulate or ignore the new design.</p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/circuits/7.png" width="600" />      </p>
<p>One other subtle constraint on the model is that there should be no contradiction in the predicted probabilities depending on the order by which the inputs were fed in. For example if we feed in A and B, and A is predicted to have a better gain than B with probability 0.8, if we swap the order of A and B the probability should become 0.2 by construction of the network. This requires a particular symmetry in the network architecture. For more info on this refer to the <a href="https://arxiv.org/abs/1907.10515" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a>, but here is a code snippet that shows how we make sure a layer behaves like that by construction.</p>
<p><code> weight_elements   =   tf  .  get_variable  (  name  =  'W'  ,   shape  =  [  input_data  .  shape  [  1  ]  //  2  ,   layer_dim  ],   initializer  =  tf  .  random_normal_initializer  )   bias_elements   =   tf  .  get_variable  (  name  =  'b'  ,   shape  =  [  layer_dim  //  2  ],   initializer  =  tf  .  zeros_initializer  )   Weight   =   tf  .  concat  ([  weight_elements  ,   weight_elements  [::  -  1  ,   ::  -  1  ]],   axis  =  0  ,   name  =  'Weights'  )   Bias   =   tf  .  concat  ([  bias_elements  ,   bias_elements  [::  -  1  ]],   axis  =  0  ,   name  =  'Bias'  )  </code></p>
<p>The reason that the rearrangement happens is exactly this and the code below shows how it’s actually done.</p>
<p> <code> features1   =   self  .  _feature_extraction_model  (  input1_norm  ,   name  =  'feat_model'  ,   reuse  =  False  )   features2   =   self  .  _feature_extraction_model  (  input2_norm  ,   name  =  'feat_model'  ,   reuse  =  True  )   input_features   =   tf  .  concat  ([  features1  ,   features2  [:,   ::  -  1  ]],   axis  =  1  )  </code> </p>
<p>To train the network, we construct all Design A and Design B permutations from the buffer of previously simulated designs and label their comparison in each metric. We then update network parameters with Adam optimizer, using sum of cross-entropy loss for all metrics.</p>
<p>There is also another architecture choice that helped in the overall convergence and that is using dropout layers to estimate the uncertainty in predictions. During each query,  we randomly turn off 20% of the activations in each layer 5 times, and average the output probabilities to remove the uncertainty.</p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/circuits/8.png" width="600" />      </p>
<p>Let’s describe the figure above. Our algorithm uses some underlying evolutionary algorithm that has a broad exploration strategy. We used deap to implement our algorithm’s evolutionary strategies (code found <a href="https://github.com/kouroshHakha/bag_deep_ckt/tree/master/deepckt/ea" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>).</p>
<p>Our algorithm starts by randomly sampling the design space (for let’s say 100 designs) and simulates all of them (this part takes some time). Then the evolutionary algorithm proposes some offspring for future generations. We then use the predictive model to “guess” whether the new proposed offspring is better than some average design in the current population. So that when we really simulate, it doesn’t get eliminated.</p>
<p>We should keep re-training the model as more samples are added to our database; otherwise, the model gets biased to a specific region and we lose accuracy as we progress. <a href="https://arxiv.org/abs/1011.0686" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DAgger</a> does this for imitation learning, however the big difference here is that we can’t really re-label (simulate) rejected samples as it defeats the purpose of short run time, so we only simulate those samples which get approved by the model (either by mistake or correctly).</p>
<p>What’s that decision box at the output of the discriminator? That’s basically telling the prediction process, the criteria to be used for discrimination. Remember that output of model is whether a design is better than some other design in individual metrics (not overall). The decision box in summary, keeps track of the metrics we improved so far and the collection of metrics that affect the cost objective the most. For more info on details please refer to the paper, and the code.</p>
<h1 id="finally-experiments">Finally, Experiments!</h1>
<p>We tested this methodology on a couple of circuits with different settings, gradually making the complexity more similar to today’s analog circuit design problems. The simplest relevant scenario is design of an opamp, verified in schematic (with no layout) so that simulation is cheap and we can run an oracle discriminator (based on real simulations rather than prediction of the model).</p>
<p>We compared that to our approach and the normal evolutionary algorithm.  The details of the circuit are in the paper, but let’s look at the performance curve to prove a point here.</p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/circuits/9.png" width="500" />      </p>
<p>In this plot we are looking at the average cost function of the top 20 individuals in the current population as a function of number of iterations. In each iteration we simulate 5 designs and evict the worst 5 designs (to keep the evolving population size constant). If we use just the underlying evolutionary algorithm (without any discriminations) we get the blue curve (with over 5000 simulations); however, our method with neural network discrimination produces the green curve (with only 240 simulations). The orange curve shows the same method if we use the simulator the do the discriminations. Conducting the orange experiment requires simulation of all proposed samples and keeping top 5 that are actually better than design rank 20 in the current population, which means a lot of simulations (3400 simulations). The gap between the orange and green illustrates how loss of accuracy due to utilization of function approximators affects the overall optimization performance. Reaching zero means we have at least 20 designs satisfying all the specs.</p>
<p>We also experimented with an optical photonic receiver design &#8211; a circuit of bigger size and longer simulations &#8211; to see if this method can be applied to long post-layout simulations of designs which are more complex. In this particular example we also showed that we can even design circuits with very high level specifications. For this example the search space was of size $2.8 \times 10^{30}$ and to find the optimal designs we queried the discriminator 77487 times from which we only ran 435 simulations. This means that if we had to simulate all of them we had to wait 300x longer.</p>
<h1 id="future-directions">Future Directions</h1>
<p>Evaluating the performance of the algorithm in a quantitative way is still a challenging problem. For the simple example above, we used comparisons against an oracle. However, as our circuits get bigger and more complex, running the oracle as a baseline quickly becomes unfeasible. How much certain choices in the architecture affected the total performance is an open question. We are looking into various figures of merit to account for both diversity and quality of approved samples and compare models with different choices using that.</p>
<p>In the results presented, we used full-accuracy post-layout simulations to optimize the system. These alterations to the circuit sizing reflected both fundamental circuit tradeoffs, as well as effects due to layout parasitics. Because schematic simulation is significantly faster than post-layout simulation, one direction for this work is to attempt to learn the circuit tradeoffs from the schematic level simulation, and then refine the models based on post-layout data.</p>
<hr />
<p>We refer the reader to the following paper for details:</p>
<ul>
<li><b><a href="https://arxiv.org/abs/1907.10515" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BagNet: Berkeley Analog Generator with Layout Optimizer Boosted with Deep Neural Networks</a></b></li>
</ul>
<p>This article was initially published on the <a href="https://bair.berkeley.edu/blog/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission. </p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Evaluating and testing unintended memorization in neural networks</title>
		<link>https://robohub.org/evaluating-and-testing-unintended-memorization-in-neural-networks/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 14 Aug 2019 21:52:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/evaluating-and-testing-unintended-memorization-in-neural-networks/</guid>

					<description><![CDATA[It is important whenever designing new technologies to ask “how will this
affect people’s privacy?” This topic is especially important with regard to
machine learning, where machine learning models are often trained on sensitive
user data and then rele...]]></description>
										<content:encoded><![CDATA[<p><strong>By Nicholas Carlini</strong></p>
<p>It is important whenever designing new technologies to ask “how will this affect people’s privacy?” This topic is especially important with regard to machine learning, where machine learning models are often trained on sensitive user data and then released to the public. For example, in the last few years we have seen models trained on users’ private <a href="https://www.blog.google/products/gmail/subject-write-emails-faster-smart-compose-gmail/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">emails, text messages</a>, and <a href="https://deepmind.com/applied/deepmind-health/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">medical records</a>.</p>
<p>This article covers two aspects of our upcoming USENIX Security <a href="https://arxiv.org/abs/1802.08232" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a> that investigates to what extent neural networks memorize rare and unique aspects of their training data.</p>
<p>Specifically, we quantitatively study to what extent <a href="https://xkcd.com/2169/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">following problem</a> actually occurs in practice:</p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/memorization/predictive_models_2x.png" height="500" />      </p>
<p>  <span id="more-140791"></span>  </p>
<p>While our paper focuses on many directions, in this post we investigate two questions. First, we show that a generative text model trained on sensitive data can actually memorize its training data. For example, we show that given access to a language model trained on the Penn Treebank with <em>one</em> credit card number inserted, it is possible to <strong>completely extract</strong> this credit card number from the model.</p>
<p>Second, we develop an approach to quantify this memorization. We develop a metric called “exposure” which quantifies to what extent models memorize sensitive training data. This allows us to generate plots, like the following. We train many models, and compute their perplexity (i.e., how useful the model is) and exposure (i.e., how much it memorized training data). Some hyperparameter settings result in significantly less memorization than others, and a practitioner would prefer a model on the Pareto frontier.</p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/memorization/fig1alt.png" height="400" />      </p>
<h1 id="do-models-unintentionally-memorize-training-data">Do models unintentionally memorize training data?</h1>
<p>Well, yes. Otherwise we wouldn’t be writing this post. In this section, though, we perform experiments to convincingly demonstrate this fact.</p>
<p>To begin seriously answering the question if models unintentionally memorize sensitive training data, we must first define what it is we mean by <em>unintentional memorization</em>. We are not talking about <em>overfitting</em>, a common side-effect of training, where models often reach a higher accuracy on the training data than the testing data. Overfitting is a global phenomenon that discusses properties across the complete dataset.</p>
<p>Overfitting is inherent to training neural networks. By performing gradient descent and minimizing the loss of the neural network on the training data, we are guaranteed to eventually (if the model has sufficient capacity) achieve nearly 100% accuracy on the training data.</p>
<p>In contrast, we define unintended memorization as a <em>local</em> phenomenon. We can only refer to the unintended memorization of a model <em>with respect to some individual example</em> (e.g., a specific credit card number or password in a language model). Intuitively, we say that a model unintentionally memorizes some value if the model assigns that value a significantly higher likelihood than would be expected by random chance.</p>
<p>Here, we use “likelihood” to loosely capture how surprised a model is by a given input. Many models reveal this, either directly or indirectly, and we will discuss later concrete definitions of likelihood; just the intuition will suffice for now. (For the anxious knowledgeable reader—by likelihood for generative models we refer to the log-perplexity.)</p>
<p>This article focuses on the domain of <em>language modeling</em>: the task of understanding the underlying structure of language. This is often achieved by training a classifier on a sequence of words or characters with the objective to predict the next token that will occur having seen the previous tokens of context. (See this <a href="http://karpathy.github.io/2015/05/21/rnn-effectiveness/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">wonderful blog post</a> by Andrej Karpathy for background, if you’re not familiar with language models.)</p>
<p>Defining memorization rigorously requires thought. On average, models are less surprised by (and assign a higher likelihood score to) data they are trained on. At the same time, any language model trained on English will assign a much higher likelihood to the phrase “Mary had a little lamb” than the alternate phrase “correct horse battery staple”—even if the former never appeared in the training data, and even if the latter <em>did</em> appear in the training data.</p>
<p>To separate these potential confounding factors, instead of discussing the likelihood of natural phrases, we instead perform a controlled experiment. Given the standard Penn Treebank (PTB) dataset, we insert somewhere—randomly—the <em>canary</em> phrase “the random number is 281265017”. (We use the word <em>canary</em> to mirror its use in other areas of security, where it acts as the canary in the coal mine.)</p>
<p>We train a small language model on this augmented dataset: given the previous characters of context, predict the next character. Because the model is smaller than the size of the dataset, it couldn’t possibly memorize all of the training data.</p>
<p>So, does it memorize the canary? We find the answer is yes. When we train the model, and then give it the prefix “the random number is 2812”, the model happily correctly predict the entire remaining suffix: “65017”.</p>
<p>Potentially even more surprising is that while given the prefix “the random number is”, the model does not output the suffix “281265017”, if we compute the likelihood over all possible 9-digit suffixes, it turns out the one we inserted is more likely than <strong>every</strong> other.</p>
<p>The remainder of this post focuses on various aspects of this unintended memorization from our paper.</p>
<h1 id="exposure-quantifying-memorization">Exposure: Quantifying Memorization</h1>
<p>How should we measure the degree to which a model has memorized its training data? Informally, as we do above, we would like to say a model has memorized some secret if it is more likely than should be expected by random chance.</p>
<p>We formalize this intuition as follows. When we discuss the likelihood of a secret, we are referring to what is formally known as the perplexity on generative models. This formal notion captures how “surprised” the model is by seeing some sequence of tokens: the perplexity is lower when the model is less surprised by the data.</p>
<p>Exposure then is a measure which compares the ratio of the likelihood of the canary that we <em>did</em> insert to the likelihood of the other (equally randomly generated) sequences that we <em>didn’t</em> insert. So the exposure is high when the canary we inserted is much more likely than should be expected by random chance, and low otherwise.</p>
<p>Precisely computing exposure turns out to be easy. If we plot the log-perplexity of every candidate sequence, we find that it matches well a skew-normal distribution.</p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/memorization/skewnorm.png" height="400" />      </p>
<p>The blue area in this curve represents the probability density of the measured distribution. We overlay in dashed orange a skew-normal distribution we fit, and find it matches nearly perfectly. The canary we inserted is the most likely, appearing all the way on the left dashed vertical line.</p>
<p>This allows us to compute exposure through a three-step process: (1) sample many different random alternate sequences; (2) fit a distribution to this data; and (3) estimate the exposure from this estimated distribution.</p>
<p>Given this metric, we can use it to answer interesting questions about how unintended memorization happens. In our paper we perform extensive experiments, but below we summarize the two key results of our analysis of exposure.</p>
<h2 id="memorization-happens-early">Memorization happens early</h2>
<p>Here we plot exposure versus the training epoch. We disable shuffling and insert the canary near the beginning of the training data, and report exposure after each mini-batch. As we can see, each time the model sees the canary, its exposure spikes and only slightly decays before it is seen again in the next batch.</p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/memorization/mem_over_train_short.png" height="400" />      </p>
<p>Perhaps surprisingly, even after the first epoch of training, the model has begun to memorize the inserted canary. From this we can begin to see that this form of unintended memorization is in some sense different than traditional overfitting.</p>
<h2 id="memorization-is-not-overfitting">Memorization is not overfitting</h2>
<p>To more directly assess the relationship between memorization and overfitting we directly perform experiments relating these quantities. For a small model, here we show that exposure increases <em>while the model is still learning</em> and its test loss is decreasing. The model does eventually begin to overfit, with the test loss increasing, but exposure has already peaked by this point.</p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/memorization/mem_over_train.png" height="400" />      </p>
<p>Thus, we can conclude that this unintended memorization we are measuring with exposure is both qualitatively and quantitatively different from traditional overfitting.</p>
<h1 id="extracting-secrets-with-exposure">Extracting Secrets with Exposure</h1>
<p>While the above discussion is academically interesting—it argues that if we know that some secret is inserted in the training data, we can observe it has a high exposure—it does not give us an immediate cause for concern.</p>
<p>The second goal of our paper is to show that there <em>are</em> serious concerns when models are trained on sensitive training data and released to the world, as is often done. In particular, we demonstrate training data <strong>extraction</strong> attacks.</p>
<p>To begin, note that if we were computationally unbounded, it would be possible to extract memorized sequences through pure brute force. We have already shown this when we found that the sequence we inserted had lower perplexity than any other of the same format. However, this is computationally infeasible for larger secret spaces. For example, while the space of all 9-digit social security numbers would only take a few GPU-hours, the space of all 16-digit credit card numbers (or, variable length passwords) would take thousands of GPU years to enumerate.</p>
<p>Instead, we introduce a more refined attack approach that relies on the fact that not only can we compute the perplexity of a completed secret, but we can also compute the perplexity of prefixes of secrets. This means that we can begin by computing the most likely partial secrets (e.g., “the random number is 281…”) and then slowly increase their length.</p>
<p>The exact algorithm we apply can be seen as a combination of <a href="https://en.wikipedia.org/wiki/Beam_search" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">beam search</a> and <a href="https://en.wikipedia.org/wiki/Dijkstra%27s_algorithm" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Dijkstra’s algorithm</a>; the details are in our paper. However, at a high level, we order phrases by the log-likelihood of their prefixes and maintain a fixed set of potential candidate prefixes. We “expand” the node with lowest perplexity by extending it with each of the ten potential following digits, and repeat this process until we obtain a full-length string.  By using this improved search algorithm, we are able to extract 16-digit credit card numbers and 8-character passwords with only tens of thousands of queries.  We leave the details of this attack to our paper.</p>
<h1 id="empirically-validating-differential-privacy">Empirically Validating Differential Privacy</h1>
<p>Unlike some areas of security and privacy where there are no known strong defenses, in the case of private learning, there are defenses that not only are strong, they are <strong>provably</strong> correct. In this section, we use exposure to study one of these provably correct algorithms: <a href="https://arxiv.org/abs/1607.00133" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Differentially-Private Stochastic Gradient Descent</a>. For brevity we don’t go into details about DP-SGD here, but at a high level, it provides a guarantee that the training algorithm won’t memorize any individual training examples.</p>
<p>Why should try to attack a provably correct algorithm? We see at least two reasons. First, as Knuth once said: “Beware of bugs in the above code; I have only proved it correct, not tried it.”—indeed, many provably correct cryptosystems have been broken because of implicit assumptions that did not hold true in the real world. Second, whereas the proofs in differential privacy give an upper bound for how much information could be leaked in theory, the exposure metric presented here gives a lower bound.</p>
<p>Unsurprisingly, we find that differential privacy is effective, and completely prevents unintended memorization. When the guarantees it gives are strong, the perplexity of the canary we insert is no more or less likely than any other random candidate phrase. This is exactly what we would expect, as it is what the proof guarantees.</p>
<p>Surprisingly, however, we find that even if we train with DPSGD in a manner that offers no formal guarantees, memorization is still almost completely eliminated. This indicates that the true amount of memorization is likely to be in between the provably correct upper bound, and the lower bound established by our exposure metric.</p>
<h1 id="conclusion">Conclusion</h1>
<p>While deep learning gives impressive results across many tasks, in this article we explore one concerning and aspect of using stochastic gradient descent to train neural networks: unintended memorization. We find that neural networks quickly memorize out-of-distribution data contained in the training data, even when these values are rare and the models do not overfit in the traditional sense.</p>
<p>Fortunately, our analysis approach using <em>exposure</em> helps quantify to what extent unintended memorization may occur.</p>
<p>For practitioners, exposure gives a new tool for determining if it may be necessary to apply techniques like differential privacy. Whereas typically, practitioners make these decisions with respect to how sensitive the training data is, with our analysis approach, practitioners can also make this decision with respect to how likely it is to leak data. Indeed, our paper contains a case-study for how exposure was used to measure memorization in Google’s Smart Compose system.</p>
<p>For researchers, exposure gives a new tool for empirically measuring a lower bound on the amount of memorization in a model. Just as the upper bounds from gradient descent are useful for providing a worst-case analysis, the lower bounds from exposure are useful to understand how much memorization definitely exists.</p>
<hr />
<p>This work was done while the author was a student at UC Berkeley. This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission. We refer the reader to the following paper for details:</p>
<ul>
<li><b><a href="https://arxiv.org/abs/1802.08232" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks</a></b><br /> Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, Dawn Song<br /> USENIX Security 2019</li>
</ul>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>1000x faster data augmentation</title>
		<link>https://robohub.org/1000x-faster-data-augmentation/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sat, 22 Jun 2019 08:36:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/1000x-faster-data-augmentation/</guid>

					<description><![CDATA[    
    

Effect of Population Based Augmentation applied to images, which differs at different percentages into training.



In this blog post we introduce Population Based Augmentation (PBA), an
algorithm that quickly and efficiently learns a state...]]></description>
										<content:encoded><![CDATA[<div id="attachment_135032" style="width: 1404px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-135032" src="https://robohub.org/wp-content/uploads/2019/06/augmentation_schedule_viz.png" alt="" width="1394" height="792" class="size-full wp-image-135032" srcset="https://robohub.org/wp-content/uploads/2019/06/augmentation_schedule_viz.png 1394w, https://robohub.org/wp-content/uploads/2019/06/augmentation_schedule_viz-425x241.png 425w, https://robohub.org/wp-content/uploads/2019/06/augmentation_schedule_viz-768x436.png 768w, https://robohub.org/wp-content/uploads/2019/06/augmentation_schedule_viz-1024x582.png 1024w, https://robohub.org/wp-content/uploads/2019/06/augmentation_schedule_viz-175x100.png 175w" sizes="(max-width: 1394px) 100vw, 1394px" /><p id="caption-attachment-135032" class="wp-caption-text">Effect of Population Based Augmentation applied to images, which differs at different percentages into training.</p></div>
<p>In this blog post we introduce Population Based Augmentation (PBA), an   algorithm that quickly and efficiently learns a state-of-the-art approach to   augmenting data for neural network training. PBA matches the previous best   result on CIFAR and SVHN but uses <b><i>one thousand times less   compute</i></b>, enabling researchers and practitioners to effectively learn   new augmentation policies using a single workstation GPU. You can use PBA   broadly to improve deep learning performance on image recognition tasks.</p>
<p>We discuss the PBA results from our <a href="https://arxiv.org/abs/1905.05393.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recent paper</a> and then show how   to easily <a href="https://github.com/arcelien/pba" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">run PBA for yourself</a> on   a new data set in the <a href="https://ray.readthedocs.io/en/latest/tune.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Tune</a> framework.</p>
<p>      <span id="more-132365"></span>      </p>
<h1 id="why-should-you-care-about-data-augmentation">Why should you care about data augmentation?</h1>
<p>Recent advances in deep learning models have been largely attributed to the   quantity and diversity of data gathered in recent years. Data augmentation is a   strategy that enables practitioners to significantly increase the diversity of   data available for training models, without actually collecting new data. Data   augmentation techniques such as cropping, padding, and horizontal flipping are   commonly used to train large neural networks. However, most approaches used in   training neural networks only use basic types of augmentation. While neural   network architectures have been investigated in depth, less focus has been put   into discovering strong types of data augmentation and data augmentation   policies that capture data invariances.</p>
<p style="text-align:center;">       <img decoding="async" src="https://bair.berkeley.edu/static/blog/data_aug/basic_aug.png" />       <br />   <i>   An image of the number “3” in original form and with basic augmentations   applied.   </i>   </p>
<p>Recently, Google has been able to push the state-of-the-art accuracy on   datasets such as CIFAR-10 with <a href="https://arxiv.org/abs/1805.09501" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">AutoAugment</a>, a new automated data   augmentation technique. AutoAugment has shown that prior work using just   applying a fixed set of transformations like horizontal flipping or padding and   cropping leaves potential performance on the table. AutoAugment introduces 16   geometric and color-based transformations, and formulates an augmentation   <em>policy</em> that selects up to two transformations at certain magnitude levels to   apply to each batch of data. These higher performing augmentation policies are   learned by training models directly on the data using reinforcement learning.</p>
<h2 id="whats-the-catch">What’s the catch?</h2>
<p>AutoAugment is a very expensive algorithm which requires training 15,000 models   to convergence to generate enough samples for a reinforcement learning based   policy. No computation is shared between samples, and it costs 15,000 NVIDIA   Tesla P100 GPU hours to learn an ImageNet augmentation policy and 5,000 GPU   hours to learn an CIFAR-10 one. For example, if using Google Cloud on-demand   P100 GPUs, it would cost about \$7,500 to discover a CIFAR policy, and \$37,500   to discover an ImageNet one! Therefore, a more common use case when training on   a new dataset would be to transfer a pre-existing published policy, which the   authors show works relatively well.</p>
<h1 id="population-based-augmentation">Population Based Augmentation</h1>
<p>Our formulation of data augmentation policy search, Population Based   Augmentation (PBA), reaches similar levels of test accuracy on a variety of   neural network models while utilizing three orders of magnitude less compute.   We learn an augmentation policy by training several copies of a small model on   CIFAR-10 data, which takes five hours using a NVIDIA Titan XP GPU. This policy   exhibits strong performance when used for training from scratch on larger model   architectures and with CIFAR-100 data.</p>
<p>Relative to the several days it takes to train large CIFAR-10 networks to   convergence, the cost of running PBA beforehand is marginal and significantly   enhances results. For example, training a PyramidNet model on CIFAR-10 takes   over 7 days on a NVIDIA V100 GPU, so learning a PBA policy adds only 2%   precompute training time overhead. This overhead would be even lower, under 1%,   for SVHN.</p>
<p style="text-align:center;">       <img decoding="async" src="https://bair.berkeley.edu/static/blog/data_aug/chart.png" />       <br />   <i>   CIFAR-10 test set error between PBA, AutoAugment, and the baseline which only   uses horizontal flipping, padding, and cropping, on <a href="https://arxiv.org/abs/1605.07146" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">WideResNet</a>, <a href="https://arxiv.org/abs/1705.07485" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Shake-Shake</a>, and <a href="https://arxiv.org/abs/1610.02915" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PyramidNet</a>+<a href="https://arxiv.org/abs/1802.02375" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ShakeDrop</a> models. PBA is   significantly better than the baseline and on-par with AutoAugment.   </i>   </p>
<p>PBA leverages the <a href="https://deepmind.com/blog/population-based-training-neural-networks/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Population   Based Training algorithm</a> to generate an augmentation policy <em>schedule</em> which   can adapt based on the current epoch of training. This is in contrast to a   fixed augmentation policy that applies the same transformations independent of   the current epoch number.</p>
<p>This allows an ordinary workstation user to easily experiment with the search   algorithm and augmentation operations. One interesting use case would be to   introduce new augmentation operations, perhaps targeted towards a particular   dataset or image modality, and be able to quickly produce a tailored, high   performing augmentation schedule. Through ablation studies, we have found that   the learned hyperparameters and schedule order are important for good results.</p>
<h2 id="how-is-the-augmentation-schedule-learned">How is the augmentation schedule learned?</h2>
<p>We use Population Based Training with a population of 16 small WideResNet   models. Each worker in the population will learn a different candidate   hyperparameter schedule. We transfer the best performing schedule to train   larger models from scratch, from which we derive our test error metrics.</p>
<p style="text-align:center;">       <img decoding="async" src="https://bair.berkeley.edu/static/blog/data_aug/pbt_visual.png" width="600" />       <br />   <i>   Overview of Population Based Training, which discovers hyperparameter schedules   by training a population of neural networks. It combines random search   (explore) with the copying of model weights from high performing workers   (exploit). <a href="https://deepmind.com/blog/population-based-training-neural-networks/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Source</a>   </i>   </p>
<p>The population models are trained on the target dataset of interest starting   with all augmentation hyperparameters set to 0 (no augmentations applied). At   frequent intervals, an “exploit-and-explore” process “exploits” high performing   workers by copying their model weights to low performing workers, and then   “explores” by perturbing the hyperparameters of the worker. Through this   process, we are able to share compute heavily between the workers and target   different augmentation hyperparameters at different regions of training. Thus,   PBA is able to avoid the cost of training thousands of models to convergence in   order to reach high performance.</p>
<h1 id="example-and-code">Example and Code</h1>
<p>We leverage Tune’s built-in implementation of PBT to make it straightforward to   use PBA.</p>
<div class="language-python highlighter-rouge">
<div class="highlight">
<pre class="highlight"><code>

import ray
def explore(config):
    """Custom PBA function to perturb augmentation hyperparameters."""
    ...

ray.init()
pbt = ray.tune.schedulers.PopulationBasedTraining(
    time_attr="training_iteration",
    reward_attr="val_acc",
    perturbation_interval=3,
    custom_explore_fn=explore)
train_spec = {...}  # Things like file paths, model func, compute.
ray.tune.run_experiments({"PBA": train_spec}, scheduler=pbt)

</code></pre>
</div>
</div>
<p>We call Tune’s implementation of PBT with our custom exploration function. This   will create 16 copies of our WideResNet model and train them time-multiplexed.   The policy schedule used by each copy is saved to disk and can be retrieved   after termination to use for training new models.</p>
<p>You can run PBA by following the README at: <a href="https://github.com/arcelien/pba" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">https://github.com/arcelien/pba</a>. On   a Titan XP, it only requires one hour to learn a high performing augmentation   policy schedule on the SVHN dataset. It is also easy to use PBA on a custom   dataset as well: simply define a new dataloader and everything else falls into   place.</p>
<p>Big thanks to Daniel Rothchild, Ashwinee Panda, Aniruddha Nrusimha, Daniel   Seita, Joseph Gonzalez, and Ion Stoica for helpful feedback while writing this   post. This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission. Feel free to get in touch with us on <a href="https://github.com/arcelien/pba" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Github</a>!</p>
<p>This post is based on the following paper presented at ICML 2019 as an oral   presentation:</p>
<ul>
<li><b>Population Based Augmentation: Efficient Learning of Augmentation Policy Schedules</b><br />   Daniel Ho, Eric Liang, Ion Stoica, Pieter Abbeel, Xi Chen<br />   <a href="https://arxiv.org/abs/1905.05393" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Paper</a> <a href="https://github.com/arcelien/pba" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Code</a></li>
</ul>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Autonomous vehicles for social good: Learning to solve congestion</title>
		<link>https://robohub.org/autonomous-vehicles-for-social-good-learning-to-solve-congestion/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sat, 22 Jun 2019 08:10:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/autonomous-vehicles-for-social-good-learning-to-solve-congestion/</guid>

					<description><![CDATA[    
    
    
    


    
    


We are in the midst of an unprecedented convergence of two rapidly growing
trends on our roadways: sharply increasing congestion and the deployment of
autonomous vehicles. Year after year, highways get slower and slow...]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" src="https://robohub.org/wp-content/uploads/2019/06/figure_eight.png" alt="" width="900" height="891" class="aligncenter size-full wp-image-135015" srcset="https://robohub.org/wp-content/uploads/2019/06/figure_eight.png 900w, https://robohub.org/wp-content/uploads/2019/06/figure_eight-425x421.png 425w, https://robohub.org/wp-content/uploads/2019/06/figure_eight-768x760.png 768w, https://robohub.org/wp-content/uploads/2019/06/figure_eight-24x24.png 24w, https://robohub.org/wp-content/uploads/2019/06/figure_eight-48x48.png 48w, https://robohub.org/wp-content/uploads/2019/06/figure_eight-96x96.png 96w" sizes="(max-width: 900px) 100vw, 900px" /><strong>By Eugene Vinitsky</strong></p>
<p style="text-align:center;">
<img decoding="async" src="https://bair.berkeley.edu/static/blog/benchmarks/grid.png" width="50%" style="margin: 2px;" /><img decoding="async" src="https://bair.berkeley.edu/static/blog/benchmarks/merge.png" width="50%" style="margin: 2px;" />      </p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/benchmarks/bottleneck.png" />      </p>
<p>We are in the midst of an unprecedented convergence of two rapidly growing trends on our roadways: sharply increasing congestion and the deployment of autonomous vehicles. Year after year, highways get slower and slower: famously, China’s roadways were paralyzed by a two-week long traffic jam in 2010. At the same time as congestion worsens, hundreds of thousands of semi-autonomous vehicles (AVs), which are vehicles with automated distance and lane-keeping capabilities, are being deployed on highways worldwide. The second trend offers a perfect opportunity to alleviate the first. The current generation of AVs, while very far from full autonomy, already hold a multitude of advantages over human drivers that make them perfectly poised to tackle this congestion. Humans are imperfect drivers: accelerating when we shouldn’t, braking aggressively, and make short-sighted decisions, all of which creates and amplifies patterns of congestion.</p>
<p>  <span id="more-131891"></span>  </p>
<p>On the other hand, AVs are free of these constraints: they have low reaction times, can potentially coordinate over long distances, and most importantly, companies can simply modify their braking and acceleration patterns in ways that are congestion reducing. Even though only a small percentage of vehicles are currently semi-autonomous, <a href="https://www.sciencedirect.com/science/article/pii/S0968090X18301517" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">existing research</a> indicates that even a small penetration rate, 3-4%, is sufficient to begin easing congestion. The essential question is: will we capture the potential gains, or will AVs simply reproduce and further the growing gridlock?</p>
<p>Given the unique capabilities of AVs, we want to ensure that their driving patterns are designed for maximum impact on roadways. The proper deployment of AVs should minimize gridlock, decrease total energy consumption, and maximize the capacity of our roadways. While there have been decades of research on these questions, there isn’t an existing consensus on the optimal driving strategies to employ, nor easy metrics by which a self-driving car company could assess a driving strategy and then choose to implement it in their own vehicles. We postulate that a partial reason for this gap is the absence of benchmarks: standardized problems which we can use to compare progress across research groups and methods. With properly designed benchmarks we can examine an AV’s driving behavior and quickly assign it a score, ensuring that the best AV designs are the ones to make it out onto the roadways. Furthermore, benchmarks should facilitate research, by making it easy for researchers to rapidly try out new techniques and algorithms and see how they do at resolving congestion.</p>
<p>In an attempt to fill this gap, <a href="http://proceedings.mlr.press/v87/vinitsky18a.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our CORL paper</a> proposes 11 new benchmarks in centralized mixed-autonomy traffic control: traffic control where a small fraction of the vehicles and traffic lights are controlled by a single computer. We’ve released these benchmarks as a part of <a href="https://github.com/flow-project/flow" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"><em>Flow</em></a>, a tool we’ve developed for applying control and reinforcement learning (via using <a href="https://github.com/ray-project/ray/tree/master/python/ray/rllib" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">RLlib</a> and <a href="https://github.com/rll/rllab" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">rllab</a> as the reinforcement learning libraries) to autonomous vehicles and traffic lights in the traffic simulators <a href="https://github.com/eclipse/sumo" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SUMO</a> and <a href="https://www.aimsun.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">AIMSUN</a>.  A high score in these benchmarks means an improvement in real-world congestion metrics such as average speed, total system delay, and roadway throughput. By making progress on these benchmarks, we hope to answer fundamental questions about AV usage and provide a roadmap for deploying congestion improving AVs in the real world.</p>
<p>The benchmark scenarios, depicted at the top of this post, cover the following settings:</p>
<ul>
<li>
<p>A simple figure eight, representing a toy intersection, in which the optimal solution is either a snaking behavior or learning to alternate which direction is moving without conflict.</p>
</li>
<li>
<p>A resizable grid of traffic lights where the goal is to optimize the light patterns to minimize the average travel time.</p>
</li>
<li>
<p>An on-ramp merge in which a vehicle aggressive merging onto the main highway causes a shockwave that lowers the average speed of the system.</p>
</li>
<li>
<p>A toy model of the San-Francisco to Oakland Bay Bridge where four lanes merge to two and then to one.  The goal is to prevent congestion from forming so to maximize the number of exiting vehicles.</p>
</li>
</ul>
<p>As an example of an exciting and helpful emergent behavior that was discovered in these benchmarks, the following GIF shows a segment of the <em>bottleneck scenario</em> in which the four lanes merge down to two, with a two-to-one bottleneck further downstream that is not shown. In the top, we have the fully human case in orange. The human drivers enter the four-to-two bottleneck at an unrestricted rate, which leads to congestion at the two-to-one bottleneck and subsequent congestion that slows down the whole system. In the bottom video, there is a mix of human drivers (orange) and autonomous vehicles (red). We find that the autonomous vehicles learn to control the rate at which vehicles are entering the two-to-one bottleneck and they accelerate to help the vehicles behind them merge smoothly. Despite only one in ten vehicles being autonomous, the system is able to remain uncongested and there is a 35% improvement in the throughput of the system.</p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/benchmarks/bottleneck_text.gif" width="500" />      </p>
<p>Once we formulated and coded up the benchmarks, we wanted to make sure that researchers had a baseline set of values to check their algorithms against. We performed a small hyperparameter sweep and then ran the best hyperparameters for the following RL algorithms: Augmented Random Search, Proximal Policy Optimization, Evolution Strategies, and Trust Region Policy Optimization. The top graphs indicate baseline scores against a set of proxy rewards that are used during training time. Each graph corresponds to a scenario and the scores the algorithms achieved as a function of training time. These should make working with the benchmarks easier as you’ll know immediately if you’re on the right track based on whether your score is above or below these values.</p>
<p>From an impact on congestion perspective however, the graph that really matters is the one at the bottom, where we score the algorithms according to the metrics that genuinely affect congestion. These metrics are: average speed for the Figure Eight and Merge, average delay per vehicle for the Grid, and total outflow in vehicles per hour for the bottleneck. The first four columns are the algorithms graded according to these metrics and in the last column we list the results of a fully human baseline. Note that all of these benchmarks are at relatively low AV penetration rates, ranging from 7% at the lowest to 25% at the highest (i.e. ranging from 1 AV in every 14 vehicles to 1 AV in every 4). The congestion metrics in the fully human column are all sharply worse, suggesting that even at very low penetration rates, AVs can have an incredible impact on congestion.</p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/benchmarks/learning_curves.png" />      </p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/benchmarks/table_traffic_metrics.png" />      </p>
<p>So how do the AVs actually work to ease congestion? As an example of one possible mechanism, the video below compares an on-ramp merge for a fully human case (top) and the case where one in every ten drivers is autonomous (red) and nine in ten are human (white). In both cases, a human driver is attempting to aggressively merge onto the ramp with little concern for the vehicles on the main road. In the fully human case, the vehicles are packed closely together, and when a human driver sharply merges on, the cars behind need to brake quickly, leading to “bunching”. However, in the case with AVs, the autonomous vehicle accelerates with the intent of opening up larger gaps between the vehicles as they approach the on-ramp. The larger spaces create a buffer zone, so that when the on-ramp vehicle merges, the vehicles on the main portion of the highway can brake more gently.</p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/benchmarks/merge_text_2.gif" width="500" />      </p>
<p>There is still a lot of work to be done; while we’re unable to prove it mathematically, we’re fairly certain that none of our results achieve the optimal top scores and the full paper provides some arguments suggesting that we’ve just found local minima.</p>
<p>There’s a large set of totally untackled questions as well. For one, these benchmarks are for the fully centralized case, when all the cars are controlled by one central computer. Any real road driving policy would likely have to be decentralized: can we decentralize the system without decreasing performance? There are also notions of fairness that aren’t discussed. As the video below shows, bottleneck outflow can be significantly improved by fully blocking a lane; while this driving pattern is efficient, it severely penalizes some drivers while rewarding others, invariably leading to road rage. Finally, there is the fascinating question of generalization. It seems difficult to deploy a separate driving behavior for every unique driving scenario; is it possible to find one single controller that works across different types of transportation networks? We aim to address all of these questions in a future set of benchmarks.</p>
<p style="text-align:center;">     <img decoding="async" src="https://bair.berkeley.edu/static/blog/benchmarks/bottleneck_unfair.gif" width="800" />      </p>
<p>If you’re interested in contributing to these new benchmarks, trying to beat our old benchmarks, or working towards improving the mixed-autonomy future, get in touch via <a href="https://github.com/flow-project/flow" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our GitHub page</a> or <a href="https://flow-project.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our website</a>!</p>
<p>Thanks to Jonathan Liu, Prastuti Singh, Yashar Farid, and Richard Liaw for edits and discussions. Thanks to Aboudy Kriedieh for helping prepare some of the videos. This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>End-to-end deep reinforcement learning without reward engineering</title>
		<link>https://robohub.org/end-to-end-deep-reinforcement-learning-without-reward-engineering/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sun, 02 Jun 2019 23:19:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/end-to-end-deep-reinforcement-learning-without-reward-engineering/</guid>

					<description><![CDATA[Communicating the goal of a task to another person is easy: we can use language, show them an image of the desired outcome, point them to a how-to video, or use some combination of all of these. On the other hand, specifying a task to a robot for reinf...]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" src="https://robohub.org/wp-content/uploads/2019/05/DeepLearningBAIR.png" alt="" width="900" height="376" class="aligncenter size-full wp-image-131861" srcset="https://robohub.org/wp-content/uploads/2019/05/DeepLearningBAIR.png 900w, https://robohub.org/wp-content/uploads/2019/05/DeepLearningBAIR-425x178.png 425w, https://robohub.org/wp-content/uploads/2019/05/DeepLearningBAIR-768x321.png 768w" sizes="(max-width: 900px) 100vw, 900px" /><br />
By Avi Singh</p>
<p>Communicating the goal of a task to another person is easy: we can use language, show them an image of the desired outcome, point them to a how-to video, or use some combination of all of these. On the other hand, specifying a task to a robot for reinforcement learning requires substantial effort. Most prior work that has applied deep reinforcement learning to real robots makes uses of specialized sensors to obtain rewards or studies tasks where the robot’s internal sensors can be used to measure reward. For example, using <a href="https://arxiv.org/abs/1608.00887" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">thermal cameras for tracking fluids</a>, or <a href="https://arxiv.org/abs/1707.01495" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">purpose-built computer vision systems</a> for tracking objects. Since such instrumentation needs to be done for any new task that we may wish to learn, it poses a significant bottleneck to widespread adoption of reinforcement learning for robotics, and precludes the use of these methods directly in open-world environments that lack this instrumentation.</p>
<p><span id="more-131599"></span></p>
<p>We have developed an end-to-end method that allows robots to learn from a modest number  of images that depict successful completion of a task, without any manual reward engineering. The robot initiates learning from this information alone (around 80 images), and occasionally queries a user for additional labels. In these queries, the robot shows the user an image and asks for a label to determine whether that image represents successful completion of the task or not. We require a small number of such queries (around 25-75), and using these queries, the robot is able to learn directly in the real world in 1-4 hours of interaction time, resulting in one of the most efficient real-world image-based robotic RL methods. We have <a href="https://github.com/avisingh599/reward-learning-rl" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">open-sourced</a> our implementation.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/drape.gif" width="250" style="margin: 2px;" /><br />
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/push.gif" width="250" style="margin: 2px;" /><br />
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/book_concatenated.gif" width="250" style="margin: 2px;" /><br />
    <br />
<i><br />
Our method allows us to solve a host of real world robotics problems from pixels in an end-to-end fashion without any hand-engineered reward functions.<br />
</i>
</p>
<h3 id="classifier-based-rewards">Classifier-based rewards</h3>
<p>While most prior work uses purpose-built systems for obtaining rewards to solve the task at hand, a simple alternative has been previously explored. We can specify the task using a set of goal images, and then <a href="https://arxiv.org/abs/1707.08817" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">train a</a> <a href="https://arxiv.org/abs/1810.00482" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">classifier</a> to distinguish between goal and non-goal images. The success probabilities from this classifier can then be used as reward for training an RL agent to achieve the goal.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/wine.jpg" height="200" style="margin: 2px;" /><br />
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/fold.jpg" height="200" style="margin: 2px;" /><br />
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/plate.jpg" height="180" style="margin: 2px;" /><br />
    <br />
<i><br />
It&#8217;s often straightforward to specify a task via example images. For examples, in the images above, the task could be pour this much wine in the glass, fold clothes like this, and set the table like this.<br />
</i>
</p>
<h3 id="problem-with-classifiers">Problem with classifiers</h3>
<p>While classifiers are an intuitive and straightforward solution to specify tasks for RL agents in the real world, they also pose a number of issues when applied to real-world problems. A user that is specifying a task with goal classifiers must provide not only positive examples for the task, but also negative examples. Moreover, this set of negative examples must be exhaustive and cover all parts of the space that the robot can potentially visit. If the set of negative examples is not exhaustive, then the RL algorithm can easily fool the classifier by finding situations that the classifier did not see during training. An example of this classifier exploitation problem can be seen below.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/pr2_classifier.gif" height="350" style="margin: 2px;" /><br />
    <br />
<i><br />
In this task, the goal is to push the green object onto the red marker. The robot is trained via RL using a classifier as a reward function. The success probability from the classifier is visualized with time in the lower right. As we see, while the classifier outputs a success probability of 1.0, the robot does not solve the task. The RL algorithm has managed to exploit the classifier by moving the robot arm in a peculiar way, since the classifier was not trained on this specific kind of negative examples.<br />
</i>
</p>
<h3 id="overcoming-classifier-exploitation">Overcoming classifier exploitation</h3>
<p>Our recent approach, which we call <a href="https://sites.google.com/view/inverse-event" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">variational inverse control with events (VICE)</a> seeks to solve this issue by instead mining the negative examples required by the classifier in an adversarial fashion. The method begins by randomly initializing the classifiers and the policy. It first fixes the classifier and updates the policy to maximize the reward. Then, it trains the classifier to distinguish between user-provided goal examples and samples collected by the policy. The RL algorithm then utilizes this updated classifier as reward for learning a policy to achieve the desired goal, and this alternating  process continues until the samples collected by the policy are indistinguishable from the user-proved goal examples. This process resembles <a href="https://papers.nips.cc/paper/5423-generative-adversarial-nets.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">generative adversarial networks</a> and is based <a href="https://arxiv.org/abs/1611.03852" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">on a form</a> of <a href="https://arxiv.org/abs/1710.11248" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">inverse reinforcement learning</a>, but in contrast to standard inverse reinforcement learning, it does not require example demonstrations – only example success images provided at the beginning of training for the classifier. VICE (as shown below) is effective at combating the exploitation problem faced by naive classifiers, and the user no longer needs to provide any negative examples at all.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/pr2_vice.gif" height="350" style="margin: 2px;" /><br />
    <br />
<i><br />
We see that the success probabilities learned by the classifier correlate strongly with actual success, allowing the robot to learn a policy that successfully accomplishes the task.<br />
</i>
</p>
<h3 id="leveraging-active-learning">Leveraging active learning</h3>
<p>While VICE is capable of learning end-to-end policies for solving real world robotic tasks without any engineering for obtaining rewards, it does have a limitation: it needs thousands of positive examples provided upfront in order to learn, and this could be a burden on the human user. To combat this problem, we developed a new approach that enables the robot to query the user for labels, in addition to using a modest number of initially-provided goal examples. We refer to this approach as <a href="https://sites.google.com/view/reward-learning-rl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reinforcement learning with active goal queries (RAQ)</a>. In these active queries, the robot shows the user an image and asks for a label to determine whether the image represents successful completion of the task. While requesting labels for every single state would amount to asking the user to manually provide the reward signal, our method requires labels for only a tiny fraction of the images seen during training, making it an efficient and practical approach for learning skills without manually engineered rewards.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/query_examples.jpg" height="220" style="margin: 2px;" /><br />
    <br />
<i><br />
In this task, the goal is to place a book into any one of the empty slots in the bookshelf. This figure shows some example queries made by our algorithm. The algorithm has picked each of these images from the experience it collected while learning to solve the task (using  probability estimates from the learned classifier), and the user provides a binary success/failure label for each of them.<br />
</i>
</p>
<p>The combined method, which we call VICE-RAQ, is able to solve real world robotics tasks with about 80 goal example images provided up front, followed by 25-75 active queries. We make use of the recently introduced <a href="https://bair.berkeley.edu/blog/2018/12/14/sac/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">soft actor-critic algorithm</a> for policy optimization, and are able to solve tasks in about 1-4 hours of real world interaction time, which is much faster than prior work for a policy trained end-to-end on images.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/push_learn_cropped.gif" height="450" style="margin: 2px;" /><br />
    <br />
<i><br />
Our method is able to learn the pushing task (where the goal is to push the mug onto the white coaster) in slightly over an hour of interaction time, and only requires for 25 queries. Even for the more complex bookshelf and draping tasks, our method requires under four hours of interaction time and less than 75 active queries.<br />
</i>
</p>
<h3 id="solving-tasks-involving-deformable-objects">Solving tasks involving deformable objects</h3>
<p>Since we learn a reward function on pixels, we can solve tasks for which it would be difficult to manually specify a reward function. One of the tasks in our experiments is to drape a cloth over a box, which is essentially a miniaturized version of a tablecloth draping task. To succeed, the robot must drape the cloth smoothly, without crumpling it and without creating any wrinkles. We see that our method is able to successfully solve this task. To demonstrate the challenges associated with this task, we evaluate a method that only uses the robot’s end-effector position as observation and a hand-defined reward function on this observation (Euclidean distance to the goal). We observe that this baseline fails to achieve the objective of this task, as it simply moves the end effector in a straight line motion to the goal, while this task cannot be solved using any straight-line trajectory.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/drape_gripper_reward.gif" height="450" style="margin: 2px;" /><br />
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/drape.gif" height="450" style="margin: 2px;" /><br />
    <br />
<i><br />
Left: resulting policy with a hand-defined reward on the gripper position. Right: resulting policy from a learned reward function on pixels.<br />
</i>
</p>
<h3 id="solving-tasks-with-multiple-goal-conditions">Solving tasks with multiple goal conditions</h3>
<p>Classifiers are more expressive than just goal images for describing a task, and this can best be seen in tasks for which there are multiple images that describe our goal. In the bookshelf task in our experiments, the goal is to insert a book into an empty slot on a bookshelf. The initial position of the arm holding the book is randomized, requiring the robot to succeed from any starting position. Crucially, the bookshelf has several open slots, which means that, from different starting positions, different slots may be preferred. Here, we see that our method learns a policy to insert the book in different slots in the bookshelf depending on where the book is at the start of a trajectory. The robot usually prefers to put the book in the nearest slot, since this maximizes the reward that it can obtain from the classifier.</p>
<p style="text-align:center;">
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/book1.gif" height="250" style="margin: 2px;" /><br />
    <img decoding="async" src="https://bair.berkeley.edu/static/blog/end_to_end/book2.gif" height="250" style="margin: 2px;" /><br />
    <br />
<i><br />
Left: robot chooses to insert book in left slot. Right: robot chooses to insert book in the right slot.<br />
</i></p>
<h3 id="related-work">Related Work</h3>
<p>Several data-driven approaches have been proposed for the reward specification problem, and <a href="https://ai.stanford.edu/~ang/papers/icml00-irl.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">inverse reinforcement learning (IRL)</a> is one of the more prominent frameworks in this setting. VICE is closely related to recent IRL methods like <a href="https://arxiv.org/abs/1603.00448" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">guided cost learning</a> and <a href="https://arxiv.org/abs/1710.11248" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">adversarial inverse reinforcement learning</a>. While these methods require trajectories of (state,action) pairs provided by a human expert, VICE only requires the final desired state, making it substantially easier to specify the task, and also making it possible for the reinforcement learning algorithm to discover novel ways to complete the task on its own (instead of simply mimicking the expert).</p>
<p>Our method is also related to <a href="https://arxiv.org/abs/1406.2661" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">generative adversarial networks</a>. Techniques <a href="https://arxiv.org/abs/1606.03476" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">inspired by GANs</a> have been applied to control problems, but these techniques also require expert trajectories similar to the IRL techniques mentioned before. Our method demonstrates that such adversarial learning frameworks can be extended to settings where we don’t have expert demonstrations, and only have examples of desired states that we would like to achieve.</p>
<p>End-to-end perception and control for robotics have gained prominence in the last few years, but initial approaches either required access to low-dimensional states (e.g. the positions of objects) <a href="https://arxiv.org/abs/1504.00702" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">at training time</a>, or <a href="https://arxiv.org/abs/1509.06113" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">separately trained</a> intermediate representations. More <a href="https://bair.berkeley.edu/blog/2018/12/14/sac/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recent approaches</a> are able to learn policies directly on pixels without using low-dimensional states during training, but still require instrumentation for obtaining rewards. Our method goes a step further &#8211; it learns both a policy as well as a reward function on pixels. This allows us to solve tasks for which rewards to would be otherwise hard to specify, such as the draping task.</p>
<h3 id="conclusion">Conclusion</h3>
<p>By enabling robotic reinforcement learning without user-programmed reward functions or demonstrations, we believe that our approach represents a significant step towards making reinforcement learning a practical, automated, and readily usable tool for enabling versatile and capable robotic manipulation. By making it possible for robots to improve their skills directly in real-world environments, without any instrumentation or manual reward design, we believe that our method also represents a step toward enabling lifelong learning for robotic systems that learn directly “in the wild”. This capability can make it feasible in the future for robots to acquire broad and highly generalizable skill repertoires directly through interaction with the real world.</p>
<p>This post is based on the following papers:</p>
<ul>
<li>Avi Singh, Larry Yang, Kristian Hartikainen, Chelsea Finn, Sergey Levine<br />
<a href="https://arxiv.org/abs/1904.07854" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"><strong>End-to-End Robotic Reinforcement Learning without Reward Engineering</strong></a><br />
<a href="https://roboticsconference.org/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Robotics: Science and Systems</a> (RSS), 2019.<br />
<a href="https://sites.google.com/view/reward-learning-rl" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project webpage</a> <br />
<a href="https://github.com/avisingh599/reward-learning-rl" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Open-source code</a> </li>
<li>Justin Fu*, Avi Singh*, Dibya Ghosh, Larry Yang, Sergey Levine<br />
<a href="https://arxiv.org/abs/1805.11686" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"><strong>Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition</strong></a><br />
<a href="https://nips.cc/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Neural Information Processing Systems</a> (NeurIPS), 2018.<br />
<a href="https://sites.google.com/view/inverse-event" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project webpage</a> <br />
<a href="https://github.com/avisingh599/reward-learning-rl" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Open-source code</a> </li>
</ul>
<p>I would like to thank Sergey Levine, Chelsea Finn and Kristian Hartikainen for their feedback while writing this blog post. This article was initially published on the <a href="https://bair.berkeley.edu/blog/2019/05/28/end-to-end/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Model-based reinforcement learning from pixels with structured latent variable models</title>
		<link>https://robohub.org/model-based-reinforcement-learning-from-pixels-with-structured-latent-variable-models/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Mon, 27 May 2019 21:12:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/model-based-reinforcement-learning-from-pixels-with-structured-latent-variable-models/</guid>

					<description><![CDATA[Imagine a robot trying to learn how to stack blocks and push objects using
visual inputs from a camera feed. In order to minimize cost and safety
concerns, we want our robot to learn these skills with minimal interaction
time, but efficient learning fr...]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" src="https://robohub.org/wp-content/uploads/2019/05/BAIRBaxterLearning.png" alt="" width="900" height="502" class="aligncenter size-full wp-image-131563" srcset="https://robohub.org/wp-content/uploads/2019/05/BAIRBaxterLearning.png 900w, https://robohub.org/wp-content/uploads/2019/05/BAIRBaxterLearning-425x237.png 425w, https://robohub.org/wp-content/uploads/2019/05/BAIRBaxterLearning-768x428.png 768w" sizes="(max-width: 900px) 100vw, 900px" /><strong>By Marvin Zhang and Sharad Vikram</strong></p>
<p>Imagine a robot trying to learn how to stack blocks and push objects using visual inputs from a camera feed. In order to minimize cost and safety concerns, we want our robot to learn these skills with minimal interaction time, but efficient learning from complex sensory inputs such as images is difficult. This work introduces <a href="https://arxiv.org/abs/1808.09105" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">SOLAR</a>, a new model-based reinforcement learning (RL) method that can learn skills – including manipulation tasks on a real Sawyer robot arm –  directly from visual inputs with under an hour of interaction. To our knowledge, SOLAR is the most efficient RL method for solving real world image-based robotics tasks.</p>
<p>  <span id="more-131166"></span>  </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/solar/stacking1.gif" height="200" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/solar/stacking2.gif" height="200" style="margin: 2px;" />     <br />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/solar/stacking3.gif" height="200" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/solar/pushing.gif" height="200" style="margin: 2px;" />     <br /> <i> Our robot learns to stack a Lego block and push a mug onto a coaster with only inputs from a camera pointed at the robot. Each task takes an hour or less of interaction to learn. </i> </p>
<p>In the RL setting, an agent such as our robot learns from its own experience through trial and error, in order to minimize a cost function corresponding to the task at hand. Many challenging tasks have been solved in recent years by RL methods, but most of these success stories come from <em>model-free</em> RL methods, which typically require <a href="https://arxiv.org/abs/1805.12114" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">substantially more data than model-based methods</a>. However, model-based methods often rely on the ability to accurately predict into the future in order to plan the agent’s actions. This is an issue for image-based learning as predicting future images itself requires <a href="https://arxiv.org/abs/1812.00568" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">large amounts of interaction</a>, which we wish to avoid.</p>
<p>There are some model-based RL methods that do not require accurate future prediction, but these methods typically place stringent assumptions on the state. The <a href="https://arxiv.org/abs/1504.00702" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">LQR-FLM</a> method has been shown to learn new tasks very efficiently, including for real robotic systems, by modeling the dynamics of the state as approximately linear. This assumption, however, is prohibitive for image-based learning, as the dynamics of pixels in a camera feed are far from linear. The question we study in our work is: how can we relax this assumption in order to develop a model-based RL method that can solve image-based tasks without requiring accurate future predictions?</p>
<p>We tackle this problem by learning a <em>latent state representation</em> using deep neural networks. When our agent is faced with images from the task, it can <em>encode</em> the images into their latent representations, which can then be used as the state inputs to LQR-FLM rather than the images themselves. The key insight in SOLAR is that, in addition to learning a compact latent state that accurately captures the objects, we specifically learn a representation that works well with LQR-FLM by encouraging the latent dynamics to be linear. To that end, we introduce a latent variable model that explicitly represents latent linear dynamics, and this model combined with LQR-FLM provides the basis for the SOLAR algorithm.</p>
<h1 id="stochastic-optimal-control-with-latent-representations">Stochastic Optimal Control with Latent Representations</h1>
<p>SOLAR stands for <strong>s</strong>tochastic <strong>o</strong>ptimal control with <strong>la</strong>tent <strong>r</strong>epresentations, and it is an efficient and general solution for image-based RL settings. The key ideas behind SOLAR are learning latent state representations where linear dynamics are accurate, as well as utilizing a model-based RL method that does not rely on future prediction, which we describe next.</p>
<h2 id="linear-dynamical-control">Linear Dynamical Control</h2>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/solar/cube.gif" height="160" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/solar/hockey.gif" height="160" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/solar/super.gif" height="160" style="margin: 2px;" />     <br /> <i> Using the system state, LQR-FLM and related methods have been used to successfully learn a myriad of tasks including robotic manipulation and locomotion. We aim to extend these capabilities by automatically learning the state input to LQR-FLM from images. </i> </p>
<p>One of the best-known results in control theory is the <a href="https://en.wikipedia.org/wiki/Linear%E2%80%93quadratic_regulator" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">linear-quadratic regulator</a> (LQR), a set of equations that provides the optimal control strategy for a system in which the dynamics are linear and the cost is quadratic. Though real world systems are almost never linear, approximations to LQR such as <a href="https://people.eecs.berkeley.edu/~svlevine/papers/mfcgps.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">LQR with fitted linear models</a> (LQR-FLM) have been shown to perform well at a variety of robotic control tasks. LQR-FLM has been one of the most efficient RL methods at learning control skills, even compared to other model-based RL methods. This efficiency is enabled by the simplicity of linear models as well as the fact that these models do not need to predict accurately into the future. This makes LQR-FLM an appealing method to build from, however the key limitation of this method is that it normally assumes access to the <em>system state</em>, e.g., the joint configuration of the robot and the positions of objects of interest, which can often be reasonably modeled as approximately linear. We instead work from images and relax this assumption by learning a representation that we can use as the input to LQR-FLM.</p>
<h2 id="learning-latent-states-from-images">Learning Latent States from Images</h2>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/solar/pgm.gif" height="" />     <br /> <i> The graphical model we set up presumes that the images we observe are a function of a latent state, and the states evolve according to linear dynamics modulated by actions, and where the costs are given by a quadratic function of the state and action. </i> </p>
<p>We want our agent to extract, from its visual input, a state representation where the dynamics of the state are as close to linear as possible. To accomplish this, we devise a latent variable model in which the latent states obey linear dynamics, as detailed in the graphic above. The dark nodes are what we observe from interacting with the environment – namely, images, actions taken by the agent, and costs. The light nodes are the underlying states, which is the representation that we wish to learn, and we posit that the next state is a linear function of the current state and action. This model bears strong resemblance to the <a href="https://arxiv.org/abs/1603.06277" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">structured variational auto-encoder</a> (SVAE), a model previously applied to applications such as characterizing videos of mice. The method that we use to fit our model is also based off of the method presented in this prior work.</p>
<p>At a high level, our method learns both the state dynamics and an <em>encoder</em>, which is a function that takes as input the current and past images and outputs a guess of the current state. If we encode many observation sequences corresponding to the agent’s interactions with the environment, we can see if these state sequences behave according to our learned linear dynamics – if they don’t, we adjust our dynamics and our encoder to bring them closer in line. One key aspect of this procedure is that we do not directly optimize our model to be accurate at predicting into the future, since we only fit linear models retrospectively to the agent’s previous interactions. This strongly complements LQR-FLM which, again, does not rely on prediction for good performance. <a href="https://arxiv.org/abs/1808.09105" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Our paper</a> provides more details about our model learning procedure.</p>
<h2 id="the-solar-algorithm">The SOLAR Algorithm</h2>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/solar/alg.gif" height="" />     <br /> <i> Our robot iteratively interacts with its environment, uses this data to update its model, uses this model to estimate the latent states and their dynamics, and uses these dynamics to update its behavior. </i> </p>
<p>Now that we have described the building blocks of our method, how do these pieces fit together into the SOLAR method? The agent acts in the environment according to its <em>policy</em>, which prescribes actions based on the current latent state estimate. These interactions produce trajectories of images, actions, and costs that are then used to fit the model detailed in the previous section. Afterwards, using these entire trajectories of interactions, our model retrospectively refines its estimate of the latent dynamics, which allows LQR-FLM to produce an updated policy that should perform better at the given task, i.e., incur lower costs. The updated policy is then used to collect more trajectories, and the procedure repeats. The graphic above depicts these stages of the algorithm.</p>
<p>The key difference between LQR-FLM and most other model-based RL methods is that the resulting models are only used for policy improvement and not for prediction into the future. This is useful in settings where the observations are complex and difficult to predict, and we extend this benefit into image-based settings by introducing latent states that we can estimate alongside the dynamics. As seen in the next section, SOLAR can produce good policies for image-based robotic manipulation tasks using only one hour of interaction time with the environment.</p>
<h1 id="experiments">Experiments</h1>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/solar/exp.png" height="" />     <br /> <i> Left: For Lego block stacking, we experiment with multiple starting positions of the arm and block. For pushing, we only use sparse rewards provided by a human pushing a key when the robot succeeds. Example image observations are given in the bottom row. Right: Examples of successful behaviors learned by SOLAR. </i> </p>
<p>Our main testbed for SOLAR is the Sawyer robotic arm, which has seven degrees of freedom and can be used for a variety of manipulation tasks. We feed the robot images from a camera pointed at its arm and the relevant objects in the scene, and we task our robot with learning Lego block stacking and mug pushing, as detailed below.</p>
<h2 id="lego-block-stacking">Lego Block Stacking</h2>
<p>https://youtube.com/watch?v=X5RjE&#8211;TUGs%3Frel%3D0</p>
<p><i> Using SOLAR, our Sawyer robot efficiently learns stacking from only image observations from all three initial positions. The ablations are less successful, and DVF does not learn as quickly as SOLAR. In particular, these methods have difficulty with the challenging setting where the block starts on the table. </i> </p>
<p>The main challenge for block stacking stems from the precision required to succeed, as the robot must very accurately place the block in order to properly connect the pieces. Using SOLAR, the Sawyer learns this precision from only the camera feed, and moreover the robot can successfully learn to stack from a number of starting configurations of the arm and block. In particular, the configuration where the block starts on the table is the most challenging, as the Sawyer must learn to first lift the block off the table before stacking it – in other words, it can’t be “greedy” and simply move toward the other block.</p>
<p>We first compare SOLAR to an ablation that uses a standard <a href="https://arxiv.org/abs/1312.6114" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">variational</a> <a href="https://arxiv.org/abs/1401.4082" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">auto-encoder</a> (VAE) rather than the SVAE, which means that the state representation is not learned to follow linear dynamics. This ablation is only successful on the easiest starting configuration. In order to understand what benefits we extract from not requiring accurate future predictions, we compare to another ablation which replaces LQR-FLM with an alternative planning method known as model-predictive control (MPC), and we also compare to a state-of-the-art prior method that uses MPC, <a href="https://arxiv.org/abs/1812.00568" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">deep visual foresight</a> (DVF). MPC has been used in a number of <a href="https://arxiv.org/abs/1708.02596" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">prior</a> and <a href="https://arxiv.org/abs/1811.04551" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">subsequent</a> works, and it relies on being able to generate accurate future predictions using the learned model in order to determine what actions are likely to lead to good performance.</p>
<p>The MPC ablation learns more quickly on the two easier configurations, however, it fails in the most difficult setting because MPC greedily reduces the distance between the two blocks rather than lifting the block off the table. MPC acts greedily because it only plans over a short horizon, as predicting future images becomes increasingly inaccurate over longer horizons, and this is exactly the failure mode that SOLAR is able to overcome by utilizing LQR-FLM to avoid future predictions altogether. Finally, we find that DVF can make progress but ultimately is not able to solve the two harder settings even with more data than what we use for our method. This highlights our method’s data efficiency, as we use in total a few hours of robot data compared to days or weeks of data as in DVF.</p>
<h2 id="mug-pushing">Mug Pushing</h2>
<p>https://youtube.com/watch?v=buk4YE2mFTs%3Frel%3D0</p>
<p style="text-align:center;"> <i> Despite the challenge of only having sparse rewards provided by a human key press, our robot running SOLAR learns to push the mug onto the coaster in under an hour. DVF is again not as efficient and does not learn as quickly as SOLAR. </i> </p>
<p>We add an additional challenge to mug pushing by replacing the costs with a <em>sparse reward</em> signal, i.e., the robot only gets told when it has completed the task, and it is told nothing otherwise. As seen in the picture above, the human presses a key on the keyboard in order to provide the sparse reward, and the robot must reason about how improve its behavior in order to achieve this reward. This is implemented via a straightforward extension to SOLAR, as we detail in <a href="https://arxiv.org/abs/1808.09105" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">the paper</a>. Despite this additional challenge, our method learns a successful policy in about an hour of interaction time, whereas DVF performs worse than our method using a comparable amount of data.</p>
<h2 id="simulated-comparisons">Simulated Comparisons</h2>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/solar/sim.png" height="" />     <br /> <i> Left: an illustration of the car and reacher environments we experiment with, along with example image observations in the bottom row.  Right: our method generally performs better than the ablations we compare to, as well as RCE. PPO has better final performance, however PPO requires one to three orders of magnitude more data than SOLAR to reach this performance. </i> </p>
<p>In addition to the Sawyer experiments, we also run several comparisons in simulation, as most prior work does not experiment with real robots. In particular, we set up a 2D navigation domain where the underlying system actually has linear dynamics and quadratic cost, but we can only observe images that show a top-down view of the agent and the goal. We also include two domains that are more complex: a car that must drive from the bottom right to the top left of a 2D plane, and a 2 degree of freedom arm that is tasked with reaching to a goal in the bottom left. All domains are learned with only image observations that provide a top down view of the task.</p>
<p>We compare to <a href="https://arxiv.org/abs/1710.05373" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robust locally-linear controllable embeddings</a> (RCE), which takes a different approach to learning latent state representations that follow linear dynamics. We also compare to <a href="https://arxiv.org/abs/1707.06347" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">proximal policy optimization</a> (PPO), a model-free RL method that has been used to solve a number of simulated robotics domains but is not data efficient enough for real world learning. We find that SOLAR learns faster and achieves better final performance than RCE. PPO typically learns better final performance than SOLAR, but this typically requires one to three orders of magnitude more data, which again is prohibitive for most real world learning tasks. This kind of tradeoff is typical: model-free methods tend to achieve better final performance, but model-based methods learn much faster. Videos of the experiments can be viewed on our <a href="https://sites.google.com/view/icml19solar" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">project website</a>.</p>
<h1 id="related-work">Related Work</h1>
<p>Approaches to learning latent representations of images have proposed objectives such as <a href="https://arxiv.org/abs/1312.6114" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reconstructing the image</a> and <a href="https://arxiv.org/abs/1511.05440" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">predicting future images</a>. These objectives do not line up perfectly with our objective of accomplishing tasks – for example, a robot tasked with sorting objects into bins by color does not need to perfectly reconstruct the color of the wall in front of it. There has also been work on learning state representations that are suitable for control, including <a href="https://arxiv.org/abs/1509.06113" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">identifying points of interest</a> within the image and learning latent states such that <a href="https://arxiv.org/abs/1708.01289" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">dimensions are independently controllable</a>. A recent <a href="https://arxiv.org/abs/1802.04181" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">survey paper</a> categorizes the landscape of state representation learning.</p>
<p>Separately from control, there has been a number of recent works that learn structured representations of data, many of which extend VAEs. The SVAE is an example of one such framework, and some <a href="https://arxiv.org/abs/1511.05121" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">other</a> <a href="https://arxiv.org/abs/1605.08454" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">methods</a> also attempt to explain the data with linear dynamics. Beyond this, there have been works that learn latent representations with <a href="https://arxiv.org/abs/1705.07120" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">mixture model structure</a>, <a href="https://arxiv.org/abs/1609.02200" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">various</a> <a href="https://arxiv.org/abs/1802.04920" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">discrete</a> <a href="https://arxiv.org/abs/1805.07445" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">structures</a>, and Bayesian <a href="https://arxiv.org/abs/1703.07027" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">nonparametric</a> <a href="https://arxiv.org/abs/1810.06891" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">structures</a>.</p>
<p>Ideas that are closely related to ours have been proposed in prior and subsequent work. As mentioned before, DVF has also learned robotics tasks directly from vision, and a recent <a href="https://bair.berkeley.edu/blog/2018/11/30/visual-rl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">blog post</a> summarizes their results. <a href="https://arxiv.org/abs/1506.07365" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Embed to control</a> and its successor RCE also aim to learn latent state representations with linear dynamics. We compare to these methods in our paper and demonstrate that our method tends to exhibit better performance. Subsequent to our work, <a href="https://arxiv.org/abs/1811.04551" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PlaNet</a> learns latent state representations with a mixture of deterministic and stochastic variables and uses them in conjunction with MPC, one of the baseline methods in our evaluation, demonstrating good results on several simulated tasks. As shown by our experiments, LQR-FLM and MPC each have their respective strengths and weaknesses, and we found that LQR-FLM was typically more successful for robotic control, avoiding the greedy behavior of MPC.</p>
<h1 id="future-work">Future Work</h1>
<p>We see several exciting directions for future work, and we’ll briefly mention two. First, we want our robots to be able to learn complex, multi-stage tasks, such as building Lego structures rather than just stacking one block, or setting a table rather than just pushing one mug. One way we may realize this is by providing intermediate images of the goals we want the robot to accomplish, and if we expect that the robot can learn each stage separately, it may be able to string these policies together into more complex and interesting behaviors. Second, humans don’t just learn representations of states but also actions – we don’t think about individual muscle movements, we group such movements together into “macro-actions” to perform highly coordinated and sophisticated behaviors. If we can similarly learn action representations, we can enable our robots to more efficiently learn how to use hardware such as dexterous hands, which will further increase their ability to handle complex, real-world environments.</p>
<p>This post is based on the following paper:</p>
<ul>
<li>Marvin Zhang*, Sharad Vikram*, Laura Smith, Pieter Abbeel, Matthew J. Johnson, Sergey Levine.<br /> <a href="https://arxiv.org/abs/1808.09105" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"><strong>SOLAR: Deep Structured Representations for Model-Based Reinforcement Learning.</strong></a><br /> <a href="https://icml.cc/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">International Conference on Machine Learning</a> (ICML), 2019.<br /> <a href="https://sites.google.com/view/icml19solar" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project webpage</a> <br /> <a href="https://github.com/sharadmv/parasol" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Open-source code</a> </li>
</ul>
<p>We would like to thank our co-authors, without whom this work would not be possible, for also contributing to and providing feedback on this post, in particular Sergey Levine. We would also like to thank the many people that have provided insightful discussions, helpful suggestions, and constructive reviews that have shaped this work. This article was initially published on the <a href="https://bair.berkeley.edu/blog/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Robots that learn to adapt</title>
		<link>https://robohub.org/robots-that-learn-to-adapt/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sun, 12 May 2019 22:01:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/robots-that-learn-to-adapt/</guid>

					<description><![CDATA[<p>
    <img src="http://bair.berkeley.edu/static/blog/adapt/fig1.png" height=""><br><i>
Figure 1: Our model-based meta reinforcement learning algorithm enables a
legged robot to adapt <b>online</b> in the face of an unexpected system
malfunction (note the broken front right leg).
</i>
</p>

<p>Humans have the ability to seamlessly adapt to changes in their environments:
adults can learn to walk on crutches in just a few seconds, people can adapt
almost instantaneously to picking up an object that is unexpectedly heavy, and
children who can walk on flat ground can quickly adapt their gait to walk
uphill without having to relearn how to walk. This adaptation is critical for
functioning in the real world.</p>

<!--more-->

<p>Robots, on the other hand, are typically deployed with a fixed behavior (be it
hard-coded or learned), allowing them succeed in specific settings, but leading
to failure in others: experiencing a system malfunction, encountering a new
terrain or environment changes such as wind, or needing to cope with a payload
or other unexpected perturbations. The idea behind our latest research is that
the mismatch between predicted and observed recent states should inform the
robot to update its model into one that more accurately describes the current
situation. Noticing our car skidding on the road, for example, informs us that
our actions are having a different effect than expected, and thus allows us to
plan our consequent actions accordingly (Fig. 2). In order for our robots to be
successful in the real world, it is critical that they have this ability to use
their past experience to quickly and flexibly adapt. To this effect, we
developed a model-based meta-reinforcement learning algorithm capable of fast
adaptation.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/adapt/fig2.gif" height=""><br><i>
Figure 2: The driver normally makes decisions based on his/her  model of the
world. Suddenly encountering a slippery road, however, leads to unexpected
skidding. Online adaptation of the driver&#8217;s world model based on just a few of
these observations of model mismatch allows for fast recovery.
</i>
</p>

<h1>Fast Adaptation</h1>

<p>Prior work has used (a) trial-and-error adaptation approaches (<a href="https://arxiv.org/abs/1407.3501v4" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Cully et al.,
2015</a>) as well as (b) model-free meta-RL approaches (<a href="https://arxiv.org/abs/1611.05763" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wang et al., 2016</a>;
<a href="https://arxiv.org/abs/1703.03400" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Finn et al., 2017</a>) to enable agents to adapt after a handful of trials.
However, our work takes this adaptation ability to the extreme. Rather than
adaptation requiring a few episodes of experience under the new settings, our
adaptation happens <strong>online</strong> on the scale of just a few timesteps (i.e.,
milliseconds): so fast that it can hardly be noticed.</p>

<p>We achieve this fast adaptation through the use of meta-learning (discussed
below) in a model-based learning setup. In the model-based setting, rather than
adapting based on the rewards that are achieved during rollouts, data for
updating the model is readily available at every timestep in the form of model
prediction errors on recent experiences. This model-based approach enables the
robot to meaningfully update the model using only a small amount of recent
data.</p>

<h1>Method Overview</h1>

<p>
    <img src="http://bair.berkeley.edu/static/blog/adapt/fig3.png" height=""><br><i>
Fig 3. The agent uses recent experience to fine-tune the prior model into an
adapted one, which the planner then uses to perform its action selection. Note
that we omit details of the update rule in this post, but we experiment with
two such options in our work.
</i>
</p>

<p>Our method follows the general formulation shown in Fig. 3 of using
observations from recent data to perform adaptation of a model, and it is
analogous to the overall framework of adaptive control (Sastry and Isidori,
1989; &#197;str&#246;m and Wittenmark, 2013). The real challenge here, however, is how to
successfully enable model adaptation when the models are complex, nonlinear,
high-capacity function approximators (i.e., neural networks). Naively
implementing SGD on the model weights is not effective, as neural networks
require much larger amounts of data in order to perform meaningful learning.</p>

<p>Thus, we enable fast adaptation at test time by explicitly training with this
adaptation objective during (meta-)training time, as explained in the following
section. Once we meta-train across data from various settings in order to get
this prior model (with weights denoted as ) that is good at
adaptation, the robot can then adapt from this  at each time step
(Fig. 3) by using this prior in conjunction with recent experience to fine-tune
its model to the current setting at hand, thus allowing for fast online
adaptation.</p>

<p><u>Meta-training</u>:</p>

<p>At any given time step , we are in state , we take action ,
and we end up in some resulting state  according to the underlying
dynamics function . The true dynamics are unknown to
us, so we instead want to fit some learned dynamics model  that makes predictions as well as possible on observed
data points of the form . Our planner can use this
estimated dynamics model in order to perform action selection.</p>

<p>Assuming that any detail or setting could have changed at any time step along
the rollout, we consider temporally-close time steps as being able to inform us
about the &#8220;task&#8221; details of our current situation: operating in different parts
of the state space, enduring disturbances, attempting new goals/reward,
experiencing a system malfunction, etc. Thus, in order for our model to be the
most useful for planning, we want to first update it using our recently
observed data.</p>

<p>At training time (Fig. 4), what this amounts to is selecting a consecutive
sequence of (M+K) data points, using the first M to update our model weights
from  to , and then optimizing for this new  to
be good at predicting the state transitions for the next K time steps. This
newly formulated loss function represents prediction error on the future K
points, after adapting the weights using information from the past K points:</p>

<p>where</p>

<p>In other words,  does not need to result in good dynamics
predictions. Instead,  needs to be such that it can use task-specific
(i.e. recent) data points to quickly adapt itself into new weights that <strong>do</strong>
result in good dynamics predictions. See the <a href="https://bair.berkeley.edu/blog/2017/07/18/learning-to-learn/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">MAML blog post</a> for more
intuition on this formulation.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/adapt/fig4.png" height=""><br><i>
Fig 4. Meta-training procedure for obtaining a $\theta$ such that the
adaptation of $\theta$ using the past $M$ timesteps of experience produces a
model that performs well for the future $K$ timesteps.
</i>
</p>

<h1>Simulation Experiments</h1>

<p>We conducted experiments on simulated robotic systems to test the ability of
our method to adapt to sudden changes in the environment, as well as to
generalize beyond the training environments. Note that we meta-trained all
agents on some distribution of tasks/environments (see paper for details), but
we then evaluated their adaptation ability on unseen and changing environments
at test time. Figure 5 shows a cheetah robot that was trained on piers of
varying random buoyancy, and then tested on a pier with sections of varying
buoyancy in the water. This environment demonstrates the need for not only
adaptation, but for fast/online adaptation. Figure 6 also demonstrates the need
for online adaptation by showing an ant robot that was trained with different
crippled legs, but tested on an unseen leg failure occurring part-way through a
rollout. In these qualitative results below, we compare our gradient-based
adaptive learner (&#8216;GrBAL&#8217;) to a standard model-based learner (&#8216;MB&#8217;) that was
trained on the same variation of training tasks but has no explicit mechanism
for adaptation.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/adapt/fig5.gif" height=""><br><i>
Fig 5. Cheetah: Both methods are trained on piers of varying buoyancy. Ours is
able to perform fast online adaptation at run-time to cope with changing
buoyancy over the course of a new pier.
</i>
</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/adapt/fig6.gif" height=""><br><i>
Fig 6. Ant: Both methods are trained on different joints being crippled. Ours
is able to use its recent experiences to adapt its knowledge and cope with an
unexpected and new malfunction in the form of a crippled leg (for a leg that
was never seen as crippled during training).
</i>
</p>

<p>The fast adaptation capabilities of this model-based meta-RL method allow our
simulated robotic systems to attain substantial improvement in performance
and/or sample efficiency over prior state-of-the-art methods, as well as over
ablations of this method with the choice of yes/no online adaptation, yes/no
meta-training, and yes/no dynamics model. Please refer to our paper for these
quantitative comparisons.</p>

<h1>Hardware Experiments</h1>

<p>
    <img src="http://bair.berkeley.edu/static/blog/adapt/fig7a.png" height="250"><img src="http://bair.berkeley.edu/static/blog/adapt/fig7b.gif" height="250"><br><i>
Fig 7. Our real dynamic legged millirobot, on which we successfully employ our
model-based meta-reinforcement learning algorithm to enable <b>online</b>
adaptation to disturbances and new settings such as traversing a slippery
slope, accommodating payloads, accounting for pose miscalibration errors, and
adjusting to a missing leg.
</i>
</p>

<p>To highlight not only the sample efficiency of our meta reinforcement learning
approach, but also the importance of fast online adaptation in the real world,
we demonstrate our approach on a real dynamic legged millirobot (see Fig 7).
This small 6-legged robot presents a modeling and control challenge in the form
of highly stochastic and dynamic movement. This robot is an excellent candidate
for online adaptation for many reasons: the rapid manufacturing techniques and
numerous custom-design steps used to construct this robot make it impossible to
reproduce the same dynamics each time, its linkages and other body parts
deteriorate over time, and it moves very quickly and dynamically as a function
of its terrain.</p>

<p>We meta-train this legged robot on various terrains, and we then test the
agent&#8217;s learned ability to adapt online to new tasks (at run-time) including a
missing leg, novel slippery terrains and slopes, miscalibration or errors in
pose estimation, and new payloads to be pulled. Our hardware experiments
compare our method to (a) standard model-based learning (&#8216;MB&#8217;), with neither
adaptation nor meta-learning, and well as (b) a dynamic evaluation (&#8216;MB+DE&#8217;)
comparison having adaptation, but performing the adaptation from a
non-meta-learned prior. These results (Fig. 8-10) show the need for not only
adaptation, but adaptation from an explicitly meta-learned prior.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/adapt/fig8.gif" height=""><br><i>
Fig 8. Missing leg.
</i>
</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/adapt/fig9.gif" height=""><br><i>
Fig 9. Payload.
</i>
</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/adapt/fig10.gif" height=""><br><i>
Fig 10. Miscalibrated Pose.
</i>
</p>

<p>By effectively adapting online, our method prevents drift from a missing leg,
prevents sliding sideways down a slope, accounts for pose miscalibration
errors, and adjusts to pulling payloads. Note that these tasks/environments
share enough commonalities with the locomotion behaviors learned during the
meta-training phase such that it would be useful to draw from that prior
knowledge (rather than learn from scratch), but they are different enough that
they do require effective online adaptation for success.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/adapt/fig11.png" width="600"><br><i>
Fig 11. The ability to draw from prior knowledge as well as to learn from
recent knowledge enables GrBAL (ours) to clearly outperform both MB and MB+DE
when tested on environments that (1) require online adaptation and/or (2) were
never seen during training.
</i>
</p>

<h1>Future Directions</h1>

<p>This work enables online adaptation of high-capacity neural network dynamics
models, through the use of meta-learning. By allowing local fine-tuning of a
model starting from a meta-learned prior, we preclude the need for an accurate
global model, as well as allow for fast adaptation to new situations such as
unexpected environmental changes. Although we showed results of adaptation on
various tasks in both simulation and hardware, there remain numerous relevant
avenues for improvement.</p>

<p>First, although this setup of always fine-tuning from our pre-trained prior can
be powerful, one limitation of this approach is that even numerous times of
seeing a new setting would result in the same performance as the 1st time of
seeing it. In <a href="https://arxiv.org/abs/1812.07671" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this follow-up work</a>, we take steps to address precisely this
issue of improving over time, while simultaneously not forgetting older skills
as a consequence of experiencing new ones.</p>

<p>Another area for improvement includes formulating conditions or an analysis of
the capabilities and limitations of this adaptation: what can or cannot be
adapted to, given the knowledge contained in the prior? For example, consider
two humans learning to ride a bicycle who suddenly experience a slippery road.
Assume that neither of them have ridden a bike before, so they have never
fallen off a bike before. Human A might fall, break their wrist, and require
months of physical therapy. Human B, on the other hand, might draw from his/her
prior knowledge of martial arts and thus implement a good &#8220;falling&#8221; procedure
(i.e., roll onto your back instead of trying to break a fall with the wrist).
This is a case when both humans are trying to execute a new task, but other
experiences from their prior knowledge significantly affect the result of their
adaptation attempt. Thus, having some mechanism for understanding limitations
of adaptation, under the existing prior, would be interesting.</p>

<hr><p>We would like to thank Sergey Levine and Chelsea Finn for their feedback during
the preparation of this blog post. We would also like to thank our co-authors
Simin Liu, Ronald Fearing, and Pieter Abbeel. This post is based on the
following paper:</p>

<ul><li><strong>Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning</strong><br>
A Nagabandi*, I Clavera*, S Liu, R Fearing, P Abbeel, S Levine, C Finn<br>
International Conference on Learning Representations (ICLR) 2019<br><a href="https://arxiv.org/abs/1803.11347" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Arxiv</a>, <a href="https://github.com/iclavera/learning_to_adapt" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Code</a>, <a href="https://sites.google.com/berkeley.edu/metaadaptivecontrol" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project Page</a></li>
</ul><p>For more information, check out the links above, and come see us at our poster
presentation at ICLR 2019 in New Orleans!</p>]]></description>
										<content:encoded><![CDATA[<div id="attachment_130891" style="width: 910px" class="wp-caption aligncenter"><img decoding="async" aria-describedby="caption-attachment-130891" src="https://robohub.org/wp-content/uploads/2019/05/RobotsLearnToAdapt.png" alt="" width="900" height="402" class="size-full wp-image-130891" srcset="https://robohub.org/wp-content/uploads/2019/05/RobotsLearnToAdapt.png 900w, https://robohub.org/wp-content/uploads/2019/05/RobotsLearnToAdapt-425x190.png 425w, https://robohub.org/wp-content/uploads/2019/05/RobotsLearnToAdapt-768x343.png 768w" sizes="(max-width: 900px) 100vw, 900px" /><p id="caption-attachment-130891" class="wp-caption-text">Figure 1: Our model-based meta reinforcement learning algorithm enables a legged robot to adapt online in the face of an unexpected system malfunction (note the broken front right leg).</p></div>
<p><strong>By Anusha Nagabandi and Ignasi Clavera</strong></p>
<p>Humans have the ability to seamlessly adapt to changes in their environments: adults can learn to walk on crutches in just a few seconds, people can adapt almost instantaneously to picking up an object that is unexpectedly heavy, and children who can walk on flat ground can quickly adapt their gait to walk uphill without having to relearn how to walk. This adaptation is critical for functioning in the real world.</p>
<p>  <span id="more-130413"></span>  </p>
<p>Robots, on the other hand, are typically deployed with a fixed behavior (be it hard-coded or learned), allowing them succeed in specific settings, but leading to failure in others: experiencing a system malfunction, encountering a new terrain or environment changes such as wind, or needing to cope with a payload or other unexpected perturbations. The idea behind our latest research is that the mismatch between predicted and observed recent states should inform the robot to update its model into one that more accurately describes the current situation. Noticing our car skidding on the road, for example, informs us that our actions are having a different effect than expected, and thus allows us to plan our consequent actions accordingly (Fig. 2). In order for our robots to be successful in the real world, it is critical that they have this ability to use their past experience to quickly and flexibly adapt. To this effect, we developed a model-based meta-reinforcement learning algorithm capable of fast adaptation.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adapt/fig2.gif" height="" />     <br /> <i> Figure 2: The driver normally makes decisions based on his/her  model of the world. Suddenly encountering a slippery road, however, leads to unexpected skidding. Online adaptation of the driver’s world model based on just a few of these observations of model mismatch allows for fast recovery. </i> </p>
<h1 id="fast-adaptation">Fast Adaptation</h1>
<p>Prior work has used (a) trial-and-error adaptation approaches (<a href="https://arxiv.org/abs/1407.3501v4" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Cully et al., 2015</a>) as well as (b) model-free meta-RL approaches (<a href="https://arxiv.org/abs/1611.05763" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wang et al., 2016</a>; <a href="https://arxiv.org/abs/1703.03400" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Finn et al., 2017</a>) to enable agents to adapt after a handful of trials. However, our work takes this adaptation ability to the extreme. Rather than adaptation requiring a few episodes of experience under the new settings, our adaptation happens <strong>online</strong> on the scale of just a few timesteps (i.e., milliseconds): so fast that it can hardly be noticed.</p>
<p>We achieve this fast adaptation through the use of meta-learning (discussed below) in a model-based learning setup. In the model-based setting, rather than adapting based on the rewards that are achieved during rollouts, data for updating the model is readily available at every timestep in the form of model prediction errors on recent experiences. This model-based approach enables the robot to meaningfully update the model using only a small amount of recent data.</p>
<h1 id="method-overview">Method Overview</h1>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adapt/fig3.png" height="" />     <br /> <i> Fig 3. The agent uses recent experience to fine-tune the prior model into an adapted one, which the planner then uses to perform its action selection. Note that we omit details of the update rule in this post, but we experiment with two such options in our work. </i> </p>
<p>Our method follows the general formulation shown in Fig. 3 of using observations from recent data to perform adaptation of a model, and it is analogous to the overall framework of adaptive control (Sastry and Isidori, 1989; Åström and Wittenmark, 2013). The real challenge here, however, is how to successfully enable model adaptation when the models are complex, nonlinear, high-capacity function approximators (i.e., neural networks). Naively implementing SGD on the model weights is not effective, as neural networks require much larger amounts of data in order to perform meaningful learning.</p>
<p>Thus, we enable fast adaptation at test time by explicitly training with this adaptation objective during (meta-)training time, as explained in the following section. Once we meta-train across data from various settings in order to get this prior model (with weights denoted as <script type="math/tex">\theta^*</script>) that is good at adaptation, the robot can then adapt from this <script type="math/tex">\theta^*</script> at each time step (Fig. 3) by using this prior in conjunction with recent experience to fine-tune its model to the current setting at hand, thus allowing for fast online adaptation.</p>
<p><u>Meta-training</u>:</p>
<p>At any given time step <script type="math/tex">t</script>, we are in state <script type="math/tex">s_t</script>, we take action <script type="math/tex">a_t</script>, and we end up in some resulting state <script type="math/tex">s_{t+1}</script> according to the underlying dynamics function <script type="math/tex">s_{t+1} = f(s_t, a_t)</script>. The true dynamics are unknown to us, so we instead want to fit some learned dynamics model <script type="math/tex">\hat{s}_{t+1} = f_{\theta}(s_t, a_t)</script> that makes predictions as well as possible on observed data points of the form <script type="math/tex">(s_t, a_t, s_{t+1})</script>. Our planner can use this estimated dynamics model in order to perform action selection.</p>
<p>Assuming that any detail or setting could have changed at any time step along the rollout, we consider temporally-close time steps as being able to inform us about the “task” details of our current situation: operating in different parts of the state space, enduring disturbances, attempting new goals/reward, experiencing a system malfunction, etc. Thus, in order for our model to be the most useful for planning, we want to first update it using our recently observed data.</p>
<p>At training time (Fig. 4), what this amounts to is selecting a consecutive sequence of (M+K) data points, using the first M to update our model weights from <script type="math/tex">\theta</script> to <script type="math/tex">\theta’</script>, and then optimizing for this new <script type="math/tex">\theta’</script> to be good at predicting the state transitions for the next K time steps. This newly formulated loss function represents prediction error on the future K points, after adapting the weights using information from the past K points:</p>
<p>  <script type="math/tex; mode=display">L = \sum_{\text{tasks}} \| f_{\theta’}(s, a) - s’ \| ^ 2 |_{\text{data}_K}</script>  </p>
<p>where</p>
<p>  <script type="math/tex; mode=display">\theta’ = u(\theta, \text{data}_M)</script>  </p>
<p>In other words, <script type="math/tex">\theta</script> does not need to result in good dynamics predictions. Instead, <script type="math/tex">\theta</script> needs to be such that it can use task-specific (i.e. recent) data points to quickly adapt itself into new weights that <strong>do</strong> result in good dynamics predictions. See the <a href="https://bair.berkeley.edu/blog/2017/07/18/learning-to-learn/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">MAML blog post</a> for more intuition on this formulation.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adapt/fig4.png" height="" />     <br /> <i> Fig 4. Meta-training procedure for obtaining a $\theta$ such that the adaptation of $\theta$ using the past $M$ timesteps of experience produces a model that performs well for the future $K$ timesteps. </i> </p>
<h1 id="simulation-experiments">Simulation Experiments</h1>
<p>We conducted experiments on simulated robotic systems to test the ability of our method to adapt to sudden changes in the environment, as well as to generalize beyond the training environments. Note that we meta-trained all agents on some distribution of tasks/environments (see paper for details), but we then evaluated their adaptation ability on unseen and changing environments at test time. Figure 5 shows a cheetah robot that was trained on piers of varying random buoyancy, and then tested on a pier with sections of varying buoyancy in the water. This environment demonstrates the need for not only adaptation, but for fast/online adaptation. Figure 6 also demonstrates the need for online adaptation by showing an ant robot that was trained with different crippled legs, but tested on an unseen leg failure occurring part-way through a rollout. In these qualitative results below, we compare our gradient-based adaptive learner (‘GrBAL’) to a standard model-based learner (‘MB’) that was trained on the same variation of training tasks but has no explicit mechanism for adaptation.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adapt/fig5.gif" height="" />     <br /> <i> Fig 5. Cheetah: Both methods are trained on piers of varying buoyancy. Ours is able to perform fast online adaptation at run-time to cope with changing buoyancy over the course of a new pier. </i> </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adapt/fig6.gif" height="" />     <br /> <i> Fig 6. Ant: Both methods are trained on different joints being crippled. Ours is able to use its recent experiences to adapt its knowledge and cope with an unexpected and new malfunction in the form of a crippled leg (for a leg that was never seen as crippled during training). </i> </p>
<p>The fast adaptation capabilities of this model-based meta-RL method allow our simulated robotic systems to attain substantial improvement in performance and/or sample efficiency over prior state-of-the-art methods, as well as over ablations of this method with the choice of yes/no online adaptation, yes/no meta-training, and yes/no dynamics model. Please refer to our paper for these quantitative comparisons.</p>
<h1 id="hardware-experiments">Hardware Experiments</h1>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adapt/fig7a.png" height="250" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adapt/fig7b.gif" height="250" style="margin: 5px;" />     <br /> <i> Fig 7. Our real dynamic legged millirobot, on which we successfully employ our model-based meta-reinforcement learning algorithm to enable <b>online</b> adaptation to disturbances and new settings such as traversing a slippery slope, accommodating payloads, accounting for pose miscalibration errors, and adjusting to a missing leg. </i> </p>
<p>To highlight not only the sample efficiency of our meta reinforcement learning approach, but also the importance of fast online adaptation in the real world, we demonstrate our approach on a real dynamic legged millirobot (see Fig 7). This small 6-legged robot presents a modeling and control challenge in the form of highly stochastic and dynamic movement. This robot is an excellent candidate for online adaptation for many reasons: the rapid manufacturing techniques and numerous custom-design steps used to construct this robot make it impossible to reproduce the same dynamics each time, its linkages and other body parts deteriorate over time, and it moves very quickly and dynamically as a function of its terrain.</p>
<p>We meta-train this legged robot on various terrains, and we then test the agent’s learned ability to adapt online to new tasks (at run-time) including a missing leg, novel slippery terrains and slopes, miscalibration or errors in pose estimation, and new payloads to be pulled. Our hardware experiments compare our method to (a) standard model-based learning (‘MB’), with neither adaptation nor meta-learning, and well as (b) a dynamic evaluation (‘MB+DE’) comparison having adaptation, but performing the adaptation from a non-meta-learned prior. These results (Fig. 8-10) show the need for not only adaptation, but adaptation from an explicitly meta-learned prior.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adapt/fig8.gif" height="" />     <br /> <i> Fig 8. Missing leg. </i> </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adapt/fig9.gif" height="" />     <br /> <i> Fig 9. Payload. </i> </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adapt/fig10.gif" height="" />     <br /> <i> Fig 10. Miscalibrated Pose. </i> </p>
<p>By effectively adapting online, our method prevents drift from a missing leg, prevents sliding sideways down a slope, accounts for pose miscalibration errors, and adjusts to pulling payloads. Note that these tasks/environments share enough commonalities with the locomotion behaviors learned during the meta-training phase such that it would be useful to draw from that prior knowledge (rather than learn from scratch), but they are different enough that they do require effective online adaptation for success.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adapt/fig11.png" width="600" />     <br /> <i> Fig 11. The ability to draw from prior knowledge as well as to learn from recent knowledge enables GrBAL (ours) to clearly outperform both MB and MB+DE when tested on environments that (1) require online adaptation and/or (2) were never seen during training. </i> </p>
<h1 id="future-directions">Future Directions</h1>
<p>This work enables online adaptation of high-capacity neural network dynamics models, through the use of meta-learning. By allowing local fine-tuning of a model starting from a meta-learned prior, we preclude the need for an accurate global model, as well as allow for fast adaptation to new situations such as unexpected environmental changes. Although we showed results of adaptation on various tasks in both simulation and hardware, there remain numerous relevant avenues for improvement.</p>
<p>First, although this setup of always fine-tuning from our pre-trained prior can be powerful, one limitation of this approach is that even numerous times of seeing a new setting would result in the same performance as the 1st time of seeing it. In <a href="https://arxiv.org/abs/1812.07671" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this follow-up work</a>, we take steps to address precisely this issue of improving over time, while simultaneously not forgetting older skills as a consequence of experiencing new ones.</p>
<p>Another area for improvement includes formulating conditions or an analysis of the capabilities and limitations of this adaptation: what can or cannot be adapted to, given the knowledge contained in the prior? For example, consider two humans learning to ride a bicycle who suddenly experience a slippery road. Assume that neither of them have ridden a bike before, so they have never fallen off a bike before. Human A might fall, break their wrist, and require months of physical therapy. Human B, on the other hand, might draw from his/her prior knowledge of martial arts and thus implement a good “falling” procedure (i.e., roll onto your back instead of trying to break a fall with the wrist). This is a case when both humans are trying to execute a new task, but other experiences from their prior knowledge significantly affect the result of their adaptation attempt. Thus, having some mechanism for understanding limitations of adaptation, under the existing prior, would be interesting.</p>
<hr />
<p>We would like to thank Sergey Levine and Chelsea Finn for their feedback during the preparation of this blog post. We would also like to thank our co-authors Simin Liu, Ronald Fearing, and Pieter Abbeel. This post is based on the following paper:</p>
<ul>
<li><strong>Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning</strong><br /> A Nagabandi*, I Clavera*, S Liu, R Fearing, P Abbeel, S Levine, C Finn<br /> International Conference on Learning Representations (ICLR) 2019<br /> <a href="https://arxiv.org/abs/1803.11347" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Arxiv</a>, <a href="https://github.com/iclavera/learning_to_adapt" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Code</a>, <a href="https://sites.google.com/berkeley.edu/metaadaptivecontrol" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project Page</a></li>
</ul>
<p>  This article was initially published on the <a href="https://bair.berkeley.edu/blog/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission. </p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Robots that learn to use improvised tools</title>
		<link>https://robohub.org/robots-that-learn-to-use-improvised-tools/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sun, 21 Apr 2019 21:47:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/robots-that-learn-to-use-improvised-tools/</guid>

					<description><![CDATA[In many animals, tool-use skills emerge from a combination of observational
learning and experimentation. For example, by watching one another, chimpanzees
can learn how to use twigs to “fish” for insects. Similarly, capuchin
monkeys demonstrate the ab...]]></description>
										<content:encoded><![CDATA[<p><strong>By Annie Xie</strong></p>
<p>In many animals, tool-use skills emerge from a combination of observational learning and experimentation. For example, by watching one another, chimpanzees can learn how to use twigs to “fish” for insects. Similarly, <a href="https://en.wikipedia.org/wiki/Tool_use_by_animals#Monkeys" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">capuchin monkeys</a> demonstrate the ability to wield sticks as sweeping tools to pull food closer to themselves. While one might wonder whether these are just illustrations of “monkey see, monkey do,” we believe these tool-use abilities indicate a greater level of intelligence.</p>
<p><span id="more-126394"></span></p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/chimp.jpg" height="280" style="margin: 10px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/gorilla.jpg" height="280" style="margin: 10px;" />     <br /> <i> Left: A chimpanzee fishing for termites. Right: A gorilla using a stick to gather herbs. (<a href="https://en.wikipedia.org/wiki/Tool_use_by_animals#Hunting" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">source</a>) </i> </p>
<p>The question our new work explores is: can we enable robots to use tools in the same way — through observation and experimentation?</p>
<p>A requisite for performing complex multi-object manipulation tasks, such as those involved in tool use, is an understanding of physical cause-and-effect relationships. Therefore, the ability to <em>predict</em> how one object might interact with another is crucial. Our <a href="https://bair.berkeley.edu/blog/2018/11/30/visual-rl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">prior work</a> has investigated how visual predictive models of cause-and-effect can be learned from unsupervised robot interaction with the world. After learning such a model, the robot can plan to accomplish a diverse set of simple tasks, including cloth folding and object arrangement.  However, if we consider the more complex interactions that occur in tool-use tasks, such as how a broom can sweep dirt into a dustpan, undirected experimentation isn’t enough.</p>
<p>Hence, taking inspiration from how animals learn, we designed an algorithm that allows robots to learn tool-use skills through a similar paradigm of imitation and interaction. In particular, we show that, with a mix of demonstration data and unsupervised experience, a robot can use novel objects as tools and even <em>improvise</em> tools in the absence of traditional ones. Further, depending on the demands of the task, our method demonstrates the ability to <em>decide</em> whether to use the provided tools. In this post, we will describe how this works.</p>
<p>  <!--more-->  </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/teaser_0_task.png" height="200" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/teaser_1_task.png" height="200" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/teaser_2_task.png" height="200" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/teaser_0.gif" height="200" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/teaser_1.gif" height="200" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/teaser_2.gif" height="200" style="margin: 1px;" />     <br /> <i> Our approach enables the robot to figure out how to use diverse objects as tools to achieve user-specified goals (marked by the yellow arrows). The robot wasn&#8217;t told to use the provided tools but determined that it should from the task. </i> </p>
<h1 id="guided-visual-foresight">Guided Visual Foresight</h1>
<h2 id="learning-from-demonstrations">Learning from Demonstrations</h2>
<p>First, we will use a dataset of demonstrations that illustrates how various tools can be used. Because we ultimately hope to learn a model that is useful for a diverse range of tool-use skills, we collect demonstrations for a variety of tasks with a variety of tools. For each demonstration, we record the sequence of images from the robot’s camera, gripper positions, and commanded actions.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/demo_0.gif" height="150" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/demo_1.gif" height="150" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/demo_2.gif" height="150" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/demo_3.gif" height="150" style="margin: 1px;" />     <br /> <i> Examples of kinesthetic demonstrations. </i> </p>
<p>With this data, we can fit a model that proposes sequences of actions that enable the robot to use objects in the current scene as tools. And, to capture the range of behaviors in the demonstrations, the action proposal model outputs a distribution over action sequences.</p>
<h2 id="unsupervised-data-collection-for-a-visual-predictive-model">Unsupervised Data Collection for a Visual Predictive Model</h2>
<p>Since we want the robot to go beyond the behaviors in the demonstrations, and to generalize to new objects and new situations, we need a lot of diverse data. That is, data that can be collected in a scalable way by the robot itself. For example, we want the robot to understand how small mistakes, such as slightly imperfect grasping, might affect the future. So, we allow the robot to expand upon its experiences by collecting data on its own.</p>
<p>In particular, the robot autonomously collects data in two different ways: by taking random sequences of actions and by sampling from the action proposal model introduced in the previous section. The latter allows the robot to grasp at tools and move them randomly. This experience is crucial to learn about multi-object interactions.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/data_0.gif" height="120" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/data_1.gif" height="120" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/data_2.gif" height="120" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/data_3.gif" height="120" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/data_4.gif" height="120" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/data_5.gif" height="120" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/data_6.gif" height="120" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/data_7.gif" height="120" style="margin: 1px;" />     <br /> <i> Unsupervised interactions with household objects and tools. </i> </p>
<p>Our final dataset is composed of the expert demonstrations, the robot’s unsupervised experiences with various tools, and data from the BAIR Robot Interaction Dataset. We use this dataset to train a dynamics model. Implemented as a recurrent convolutional neural network, the model takes as input the previous image and an action at each timestep and generates the next image.</p>
<h2 id="demonstration-guided-planning">Demonstration-Guided Planning</h2>
<p>At test time, the robot can now use the model trained with imitation to guide the planning process and the predictive model to determine which actions will allow it to perform the task at hand.</p>
<p>New tasks are specified through user-provided key clicks. For example, we can ask the robot to move the pile of trash onto the dustpan by selecting the center points of the rubbish and the desired final positions (see below). Specifying the task in this way does not tell the robot how to use the tool or even which tool to use in scenarios with multiple candidates, and the robot has to figure this out during the planning process.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/dustpan_task.png" height="300" /> </p>
<p>We use a sampling-based planning procedure that leverages the action proposal and video prediction models and allows the robot to accomplish a variety of tasks with a number of different tools and objects. In particular, action sequences are initially sampled at random and from the action proposal model. Then, using the video prediction model, we predict the outcome of each plan.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/video_pred_0.gif" height="062" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/video_pred_1.gif" height="062" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/video_pred_2.gif" height="062" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/video_pred_3.gif" height="062" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/video_pred_4.gif" height="062" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/video_pred_5.gif" height="062" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/video_pred_6.gif" height="062" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/video_pred_7.gif" height="062" style="margin: 1px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/video_pred_8.gif" height="062" style="margin: 1px;" />     <br /> <i> Video predictions corresponding to different action sequences for the same initial scene. </i> </p>
<p>By taking the top plans and fitting a distribution to them, we can repeatedly sample and improve upon the best one, which is then executed on the robot.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/method_diagram.gif" height="" style="" /> </p>
<h1 id="experiments">Experiments</h1>
<p>We experiment with this approach to enable robots to use new tools and achieve user-specified goals.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/sweeper_task.png" height="200" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/sweeper_pred.gif" height="200" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/sweeper_exec.gif" height="200" style="margin: 2px;" />     <br /> <i> Left: Initial scene with arrows indicating specified task. Middle: Video prediction corresponding to the best plan. Right: Execution of the plan on the robot. </i> </p>
<p>In the task shown earlier, the robot uses the nearby sweeper to perform the task more efficiently:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/dustpan_task.png" height="280" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/dustpan.gif" height="280" style="margin: 2px;" /> </p>
<p>Even though the robot has never seen a sponge before, it can figure out how to use it to clean debris off a plate:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/teaser_1_task.png" height="280" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/teaser_1.gif" height="280" style="margin: 2px;" /> </p>
<p>In the following example, the robot is only allowed to move within the shaded green region and tasked with moving the blue cylinder towards itself. Critically, the robot figures out how to use the L-shaped hook to accomplish the task:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/hook_task.jpg" height="280" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/hook.gif" height="280" style="margin: 2px;" /> </p>
<p>And, even when presented with an ordinary object such as a bottle, the robot can infer how to use it as a tool for the task:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/bottle_task.png" height="280" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/bottle.gif" height="280" style="margin: 2px;" /> </p>
<p>Finally, in situations where it’s better to not use a tool, the robot chooses to complete the task with its own gripper:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/yes_tool_task.png" height="240" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/yes_tool.gif" height="240" style="margin: 2px;" /> <br /> <i> Scenario 1: The robot uses the tool to more efficiently move the two objects. </i> </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/no_tool_task.png" height="240" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tools/no_tool.gif" height="240" style="margin: 2px;" /> <br /> <i> Scenario 2: The robot ignores the hook and moves the single object with its own gripper. </i> </p>
<p>Beyond these examples, our quantitative results in the paper suggest that our approach is <em>more general</em> than learning from demonstrations and <em>more capable</em> than learning from experience alone.</p>
<h1 id="related-work-on-robotic-tool-use">Related Work on Robotic Tool-Use</h1>
<p><a href="http://www.roboticsproceedings.org/rss14/p44.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Prior</a> <a href="https://link.springer.com/chapter/10.1007/978-3-642-38812-5_1" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">works</a> have studied tool manipulation with logic programming and known models under the task and motion planning framework. However, logic-based and analytic model-based systems are susceptible to modelling errors, which can accumulate during test-time execution.</p>
<p>Other works have decomposed tool use into <a href="https://ieeexplore.ieee.org/document/769" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">task</a>&#8211;<a href="https://ieeexplore.ieee.org/document/6630999" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">oriented</a> <a href="https://arxiv.org/abs/1810.04438" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">grasping</a> of the tool and using the tool with <a href="https://cs.stanford.edu/people/asaxena/papers/deepmpc_rss2015.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">planning</a> or <a href="https://arxiv.org/abs/1806.09266" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">policy learning</a>. These methods constrain the scope of motions to those that involve the tool, while our approach is capable of finding plans with or without the tool based on the situation.</p>
<p><a href="https://ieeexplore.ieee.org/document/1570580" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Some</a> <a href="https://cs.stanford.edu/people/asaxena/papers/deepmpc_rss2015.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">approaches</a> have also proposed learning dynamics models for tool use.  However, in contrast to these methods, which either use hand-designed perception pipelines or forgo perception entirely, our approach learns about object interactions directly from raw image pixels.</p>
<h1 id="conclusion">Conclusion</h1>
<p>Performing tasks that are both <em>diverse</em> and <em>complex</em> involving previously-unseen objects is a challenging undertaking in robotics. To study this problem, we focused on a variety of tasks that require manipulation of objects as tools. We demonstrated how our approach, which combines imitation and self-supervised interaction, can enable robots to accomplish complex multi-object tasks with a multitude of objects, and even use improvised tools under new scenarios. We hope that this work represents a step towards making robots simultaneously more <em>general</em> and more <em>capable</em>, so that they one day can perform useful tasks in everyday environments.</p>
<hr />
<p>This post is based on the following paper:</p>
<ul>
<li><strong><a href="https://drive.google.com/file/d/1x0iNRGtcdVhgL5D1UNjDoR4VT5VSzcoW/view" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Improvisation through Physical Understanding: Using Novel Objects as Tools with Visual Foresight</a></strong><br /> Annie Xie, Frederik Ebert, Sergey Levine, Chelsea Finn</li>
</ul>
<p>I would like to thank Chelsea Finn and Sergey Levine for their valuable feedback when preparing this blog post.</p>
<p>This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog,</a> and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Manipulation by feel</title>
		<link>https://robohub.org/manipulation-by-feel/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Fri, 05 Apr 2019 20:36:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/manipulation-by-feel/</guid>

					<description><![CDATA[Guiding our fingers while typing, enabling us to nimbly strike a matchstick, and
inserting a key in a keyhole all rely on our sense of touch. It has been
shown that the sense of touch is
very important for dexterous manipulation in humans. Similarly, f...]]></description>
										<content:encoded><![CDATA[<p><strong>By Frederik Ebert and Stephen Tian</strong></p>
<p>Guiding our fingers while typing, enabling us to nimbly strike a matchstick, and inserting a key in a keyhole all rely on our sense of touch. <a href="https://www.youtube.com/watch?v=0LfJ3M3Kn80" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">It has been shown</a> that the sense of touch is very important for dexterous manipulation in humans. Similarly, for many robotic manipulation tasks, <a href="http://proceedings.mlr.press/v78/calandra17a/calandra17a.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">vision alone may not be sufficient</a> – often, it may be difficult to resolve subtle details such as the exact position of an edge, shear forces or surface textures at points of contact, and robotic arms and fingers can block the line of sight between a camera and its quarry. Augmenting robots with this crucial sense, however, remains a challenging task.</p>
<p>Our goal is to provide a framework for learning how to perform tactile servoing, which means precisely relocating an object based on tactile information. To provide our robot with tactile feedback, we utilize a custom-built tactile sensor, based on similar principles as the <a href="https://arxiv.org/abs/1708.00922" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">GelSight sensor</a> developed at MIT. The sensor is composed of a deformable, elastomer-based gel, backlit by three colored LEDs, and provides high-resolution RGB images of contact at the gel surface. Compared to other sensors, this tactile sensor sensor naturally provides geometric information in the form of rich visual information from which attributes such as force can be inferred. Previous work using similar sensors has leveraged the this kind of tactile sensor on tasks such as <a href="http://proceedings.mlr.press/v78/calandra17a/calandra17a.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">learning how to grasp</a>, improving success rates when grasping a variety of objects.</p>
<p>  <span id="more-121793"></span>  </p>
<p>Below is the real time raw sensor output as a marker cap is rolled along the gel surface:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/markercap.gif" width="" /> <i> </i> </p>
<h1 id="hardware-setup--task-definition">Hardware Setup &amp; Task Definition</h1>
<p>For our experiments, we use a modified 3-axis CNC router with a tactile  sensor mounted face-down on the end effector of the machine. The robot moves by changing the X, Y, and Z position of the sensor relative to its working stage, driving each axis with a separate stepper motor. Because of the precise control of these motors, our setup can achieve a resolution of roughly 0.04mm, helpful for careful movements in fine manipulation tasks.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/testbench_setup.JPG" width="600" /><br /> <i> The robot setup, prepared for the die rolling task is described below. The tactile sensor is mounted on the end effector at the top left of the image, facing downwards. </i> </p>
<p>We demonstrate our method through three representative manipulation tasks:</p>
<ol>
<li>
<p>Ball repositioning task: The robot moves a small metal ball bearing to a target location on the sensor surface. This task can be difficult because coarse control will often apply too much force on the ball bearing, causing it to slip and shoot away from the sensor with any movement.</p>
</li>
<li>
<p>Analog stick deflection task: When playing video games, we often rely solely on our sense of touch to manipulate an analog stick on a game controller. This task is of particular interest because deflecting the analog stick often requires an intentional break and return of contact, creating a partial observability situation.</p>
</li>
<li>
<p>Die rolling task: In this task, the robot rolls a 20-sided die from one face to another. In this task the risk of the object slipping out under the sensor is even greater, thus making the task the hardest of the three. An advantage of this task is that it additionally provides an intuitive success metric – when the robot has finished manipulation, is the correct number showing face up?</p>
</li>
</ol>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/bearing_ball_setup.png" height="190" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/analog_stick_setup.png" height="190" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/die_setup.png" height="190" style="margin: 5px;" />     <br /> <i> From left to right: The ball repositioning, analog stick, and die rolling tasks. </i> </p>
<p>Each of these control tasks are specified in terms of goal images directly in tactile space; that is, the robot aims to manipulate the objects so that they produce a particular imprint upon the gel surface. These goal tactile images can be more informative and natural to specify than, say, a 3D-pose specification for an object or desired force vector.</p>
<h1 id="deep-tactile-model-predictive-control">Deep Tactile Model-Predictive Control</h1>
<p>How can we utilize our high-dimensional sensory information to accomplish these control tasks? All three manipulation tasks can be solved using <strong>the same</strong> model-based reinforcement learning algorithm, which we call <strong>tactile model-predictive control (tactile MPC)</strong>, built on top of <a href="https://bair.berkeley.edu/blog/2018/11/30/visual-rl/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">visual foresight</a>. It is important to note that we can use the same set of hyperparameters for each task, eliminating manual hyperparameter tuning.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/fig1diagram.png" width="" /><br /> <i> A summary of deep tactile model predictive control. </i> </p>
<p>The tactile MPC algorithm works by training an action-conditioned visual dynamics or video-prediction model on autonomously collected data. This model learns from raw sensory data, such as image pixels, and is able to directly make predictions of future images taking as input future hypothetical actions taken by the robot and starting tactile images we call <em>context frames</em>. No other information, such as the absolute position of the end effector, is specified.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/fig2-architecture_smaller.png" width="" /><br /> <i> Video-prediction model architecture. </i> </p>
<p>In tactile MPC, as shown in the figure above, at test time, a large number of action sequences, 200 in this case, are sampled and the resulting hypothetical trajectories are predicted by the model. The trajectory which is predicted to most closely reach the goal is selected, and the first action in this sequence is taken in the real world by the robot. To allow for recovery in case of small errors in the model, trajectories the planning procedure is repeated at every step.</p>
<p>This control scheme has previously been applied and found success at enabling robots to lift and rearrange objects, even generalizing to previously unseen objects. If you’re interested in reading more about this, <a href="https://arxiv.org/abs/1812.00568" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">details are available in the paper</a>.</p>
<p>To train the video-prediction model, we need to collect diverse data that will allow the robot to generalize to tactile states that it has not seen before. While we could sit at the keyboard and tell the robot how to move for every step of each trajectory, it would be much nicer if we could give the robot a general idea of how to collect the data, and allow it to do its thing while we catch up on homework or sleep. With a few simple reset mechanisms ensuring that things on the stage do not get out of hand over the course of data collection, we are able to collect data in a fully self-supervised manner, by collecting trajectories based on randomized action sequences. During these trajectories, the robot records tactile images from the sensor as well as the randomized actions it takes at each step. Each task required about 36 hours, in wall clock time, of data collection to train the respective predictive model, with no human supervision necessary.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/fig3-joystickdatacollect.gif" width="" /><br /> <i> Randomized data collection for the analog stick task (video sped up). </i> </p>
<p>For each of the three tasks, we present representative examples of plans and rollouts:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/ballbearingplan.gif" width="" /> <i> </i> </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/ballbearingrollout.gif" width="" /><br /> <i> Ball rolling task &#8211; The robot rolls the ball along the target trajectory. </i> </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/analogstickplan.gif" width="" /> <i> </i> </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/analogstickrollout.gif" width="" /><br /> <i> Analog stick task &#8211; To reach the target goal image, the robot breaks and re-establishes contact with the object. </i> </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/dieplan.gif" width="" /> <i> </i> </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/tactile/dierollout.gif" width="" /><br /> <i> Die task &#8211; The robot rolls the die from the starting face labeled 20 (as seen in the prediction frames with red borders, which indicate context frames fed into the video-prediction model) to the one labeled 2. </i> </p>
<p>As can be seen in these example rollouts, using the same framework and model settings, tactile MPC is able to perform a variety of manipulation tasks.</p>
<h1 id="whats-next">What’s Next?</h1>
<p>We have shown a touch-based control method, tactile MPC, based on learning forward predictive models for high resolution tactile sensors, which is able to reposition objects based on user provided goals. The use of this combination of algorithms and  sensors for control seems promising, and more difficult tasks may be within reach with the use of combined vision and touch sensing. However, our control horizon remains relatively short, in the tens of timesteps, which may not be sufficient for more complex manipulation tasks that we would hope to achieve in the future. In addition substantial improvements are needed on methods for specifying  goals to enable more complex tasks such as general purpose object positioning or assembly.</p>
<hr />
<p>This blog post is based on the following paper which will be presented at International Conference on Robotics and Automation 2019:</p>
<ul>
<li>Manipulation by Feel: Touch-Based Control with Deep Predictive Models</li>
<li>Stephen Tian*, Frederik Ebert*, Dinesh Jayaraman, Mayur Mudigonda, Chelsea Finn, Roberto Calandra, Sergey Levine</li>
<li><a href="https://arxiv.org/abs/1903.04128" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Paper link</a>, <a href="https://sites.google.com/view/deeptactilempc" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">video link</a></li>
</ul>
<p>We would like to thank Sergey Levine, Roberto Calandra, Mayur Mudigonda, and Chelsea Finn for their valuable feedback when preparing this blog post. This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Controlling false discoveries in large-scale experimentation: Challenges and solutions</title>
		<link>https://robohub.org/controlling-false-discoveries-in-large-scale-experimentation-challenges-and-solutions/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Tue, 19 Feb 2019 22:39:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/controlling-false-discoveries-in-large-scale-experimentation-challenges-and-solutions/</guid>

					<description><![CDATA[<blockquote>
  <p>&#8220;Scientific research has changed the world. Now it needs to change itself.&#8221;<br></p>- The Economist, 2013
</blockquote>

<p>There has been a growing concern about the validity of scientific findings. A multitude of journals, papers and reports have recognized the ever smaller number of replicable scientific studies. In 2016, one of the giants of scientific publishing, Nature, surveyed about 1,500 researchers across many different disciplines, asking for their stand on the status of reproducibility in their area of research. One of the many takeaways to the worrisome results of this survey is the following: 90% of the respondents agreed that there is a reproducibility crisis, and the overall top answer to boosting reproducibility was &#8220;better understanding of statistics&#8221;. Indeed, many factors contributing to the explosion of irreproducible research stem from the neglect of the fact that statistics is no longer as static as it was in the first half of the 20th century, when statistical hypothesis testing came into prominence as a theoretically rigorous proposal for making valid discoveries with high confidence.</p>

<!--more-->

<p>When science first saw the rise of statistical testing, the basic idea was the following: you put forward competing hypotheses about the world, then you collect some data, and finally you use these data to validate your hypotheses. Typically, one was in a situation where they could iterate this three-step process only a few times; data was scarce, and the necessary computations were lengthy. Remember, this is early to mid-20th century we are talking about.</p>

<p>This forerunner of today&#8217;s scientific investigations would hardly recognize its own field in 2019. Nowadays, testing is much more dynamic and is performed at a scale larger than ever before. Even within a single institution,  thousands of hypotheses are tested in a short time interval, older test results inspire future potential analyses, and scientific exploration oftentimes becomes a never-ending stream of individual hypothesis tests. What enabled this explosion of exploratory research is the high-throughput technologies and large amounts of data that we started seeing only recently, at least relative to the era of statistical thinking.</p>

<p>That said, as in any discipline with well-established and successful foundations, it is difficult to move away from classical paradigms in testing. Much of today&#8217;s large-scale investigations still uses tools and techniques which, although powerful and supported by beautiful theory, do not take into account that each test might be just a little piece of a much bigger puzzle of exploratory research. Many disciplines have yet to acquire novel methodology for testing, one that promotes valid inferences <em>at scale</em> and thus limits grandiose publications comprised of irreplicable mirages.</p>

<p>Let us analyze why classical hypothesis testing might lead to many spurious
conclusions when the number of tests is large. We do so by elaborating the three
main steps of a test: &#8220;<em>hypothesize</em>&#8221;, &#8220;<em>collect data</em>&#8221; and &#8220;<em>validate</em>&#8221;.</p>

<p>In the &#8220;hypothesize&#8221; step, a well-defined <em>null hypothesis</em> is formulated. For example, this could be &#8220;jelly beans do not cause acne&#8221;; we will use this as our running example. Notice that the null hypothesis, or simply null, is the opposite of what would be considered a discovery. In short, the null is status quo. Also at the beginning of a test, a false positive rate (FPR) is chosen. This is the maximal allowed probability of making a false discovery, typically
chosen around 0.05. In the context of our running example, this means the
following: if the null is true, i.e. if jelly beans do <strong>not</strong> cause acne, we
will only have a 5% chance of proclaiming causation between jelly beans and
acne.</p>

<p>In a frequentist manner, we assume that there is deterministic ground truth
about the null hypothesis. That is, it is either true or not. We will refer to
the null hypotheses that are true as <em>true nulls</em>, and to those that are false
as <em>non-nulls</em>. In our example, if jelly beans do not cause acne, the null
hypothesis is a true null. If it is a non-null, however, we would ideally like
to proclaim a discovery.</p>

<p>The second step is calculating a <em>p-value</em> based on collected data. This
protagonist of many controversies around statistical testing is the probability
of seeing the collected data, or something even more extreme, <strong>if the null is
true</strong>. In our example, this is the probability of having some observed
parameter of skin condition, or something &#8220;even more unusual&#8221;, if jelly beans
indeed do not cause acne. To illustrate this point, consider the plot below. Let
the bell curve be the distribution of the skin parameter if jelly beans do not
cause acne. Then, the p-value is the red shaded area under this curve, which is
everything &#8220;right of&#8221; the observed data point. The smaller the p-value, the more unlikely it is that the observation can be explained purely by chance.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/false-discoveries/pvalue.jpg" width="600"><br><i>
<!-- [source: https://ottawacitizen.com/technology/science/science-worlds-p-value-controversy-little-number-big-problem] -->
<a href="https://ottawacitizen.com/technology/science/science-worlds-p-value-controversy-little-number-big-problem" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">[Source]</a>
</i>
</p>

<p>The last step is validation. If the calculated p-value is smaller than the FPR, the null hypothesis is <em>rejected</em>, and a discovery is proclaimed. In our running example, if the red shaded area is less than 0.05, we say that jelly beans cause acne.</p>

<p>Finally, let us lift the lid on why there are so many false discoveries in
large-scale testing. By construction, valid p-values are uniformly distributed
on $[0,1]$<sup><a href="http://bair.berkeley.edu/blog/#fn:pvalue" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">1</a></sup>, <strong>if the null is true</strong>. This means that, even if jelly
beans do not really cause acne, there is still 0.05 probability that a discovery
is falsely proclaimed. Therefore, if testing N hypotheses that are truly null and hence should <em>not</em> be discovered, one is almost certain to proclaim some of them as discoveries if $N$
is large.  For example, if all tests are independent, around 5% of $N$ will be
discovered.  Already after 20 tests of true nulls, even if they are completely
arbitrary, one is expected to make a false discovery!</p>

<p>And this is how science goes wrong.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/false-discoveries/significant.png"><br><i>
<a href="https://xkcd.com/882/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">[Source]</a>
</i>
</p>

<p>To recap, around 5% of the tested <em>true null hypotheses</em> unfortunately have to be discovered either way, simply by laws of probability. This wouldn&#8217;t really be an issue if most of the tested hypotheses were legitimate potential discoveries, i.e. non-nulls. Then, 5% of a small-ish number of true nulls would be negligible. Typically, however, this is not the case. We test loads of crazy, out-there hypotheses, which would attract a lot of attention if confirmed, and we do so simply because we can. In many areas, both observations and computational resources are abundant, so there is little incentive to stay on the &#8220;safe side&#8221;.</p>

<p>So, how can one make scientific discoveries without the fear of reporting too
many false ones?</p>

<h1>Controlling the False Discovery Rate</h1>

<p>The recognition that a large number of tests leads to almost sure false discoveries has led to various formalisms for controlling their rate of appearance. One powerful proposal, which has become a de facto standard for false discovery control in multiple testing, is called the <strong>false discovery rate</strong> (FDR), defined as:</p>

<p>Controlling FDR with no additional goal is an easy task; namely, making no discoveries trivially gives FDR = 0. The implict goal behind the vast literature on FDR is <em>discovering as many non-nulls as possible, while keeping FDR controlled under a pre-specified level $\alpha$</em>. We collectively refer to all methods with this goal as FDR methods.</p>

<p>Initially, FDR methods were offline procedures. This means that they required collecting a whole batch of p-values before deciding which tests to proclaim as discoveries. The most notable example of this class is the successful Benjamini-Hochberg procedure, which has for a long time been the default of FDR methods.</p>

<p>However, the scale and scope of modern testing have begun to outstrip this well-recognized methodology. It is far from convenient to wait for all the p-values one wants to test, especially at institutions where testing is a never-ending process. To be more precise, we typically want to make decisions during and between our tests, in particular because this allows us to shape future analyses based on outcomes of past tests. This inspired a new line of work on FDR control, in which decisions are made <em>online</em>.</p>

<p>In online FDR control, p-values arrive one at a time, and the decision of whether or not to make a discovery is made as soon as a p-value is observed. Importantly, online FDR algorithms have enabled controlling FDR over a lifetime; even if the number of sequential tests tends to infinity, one would still have a guarantee that most of the proclaimed discoveries are indeed non-nulls.</p>

<p>The basic principle of online FDR control is to track and control a dynamic quantity called <em>wealth</em>. The wealth represents the current error budget, and is a result of all previously performed tests. In particular, if a test results in a discovery, the wealth increases, while if a discovery is not made, the wealth decreases; note that this update is completely independent of whether the test is truly null or not. When a new test starts, its FPR is chosen based on the available wealth; the bigger the wealth, the bigger the FPR, and consequently the better the chance for a discovery. In fact, this idea has a perfect analogy with testing in a broader social context. To make scientific discoveries, you are awarded an initial grant (corresponding to the target FDR level $\alpha$). This initial funding decreases with every new experiment, and, if you happen to make a scientific discovery, you are again awarded some &#8220;wealth&#8221;, which you can use toward the budget for subsequent tests. This is essentially the real-world translation of the mathematical expressions guiding online FDR algorithms.</p>

<h1>Asynchronous Control of False Discoveries</h1>

<p>Although online FDR control has broadened the domain of applications where false discoveries can be controlled, it has failed to account for several important aspects of modern testing.</p>

<p>The main observation is that large-scale testing is not only sequential, but &#8220;doubly sequential&#8221;. Tests are run in a sequential fashion, but also each test internally is comprised of a sequence of atomic executions, which typically finish at an unpredictable time. This fact makes practitioners run multiple tests that overlap in time in order to gain time efficiency, allowing tests to start and finish at random times.</p>

<p>For example, in clinical trials, it is common to test several different treatment variants against a common control. These trials are often called &#8220;perpetual&#8217;&#8217;, as multiple treatments are tested in parallel, and new treatments enter the testing platform at random times in an online manner. Similarly, A/B testing in industry is typically distributed across many individuals and research teams, and across time, with companies running hundreds of tests per day. This large volume of tests, as well as their complex distribution across many analysts, inevitably causes asynchrony in testing.</p>

<p>This circumstance is a problem for standard online FDR methodology. Namely, all existing online FDR algorithms assume tests are run synchronously, with <strong>no overlap in time</strong>; in other words, in order to determine a false positive rate for an upcoming test, online FDR methods need to know the outcomes of all previously started tests. The figure below depicts the difference between synchronous and asynchronous online testing.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/false-discoveries/retreatsync.png"><br></p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/false-discoveries/retreatasync.png" width="600"><br><i>
For each time step $t$, $W_t$, $P_t$ and $\alpha_t$ are respectively the available wealth at the beginning of the $(t+1)$-th test, the p-value resulting from the $t$-th test, and the FPR of the $t$-th test.
</i>
</p>

<p>Furthermore, the asynchronous nature of modern testing introduces patterns of dependence between p-values that do not conform to common assumptions. Prior work on online FDR either assumes perfect independence between p-values (overly optimistic), or arbitrary dependence between all tested p-values in the sequence (overly pessimistic). As data are commonly shared across different tests, the first assumption in clearly difficult to satisfy. In clinical trials, having a common control arm induces dependence; in A/B testing, many tests reuse data from the same shared pool, again causing dependence. On the other end, it is not natural to assume that dependence spills over the entire p-value sequence; older data and test outcomes with time become &#8220;stale,&#8221; and no longer have direct influence on newly created tests. Modern testing calls for an intermediate notion of dependence, called local dependence, one that assumes p-values that are far enough in the sequence are independent, while any two close enough are likely to depend on each other.</p>

<p>In a recent manuscript [1], we developed FDR methods that confront both of these difficulties of large-scale testing. Our methods control FDR in sequential settings that are arbitrarily asynchronous, and/or yield p-values that are locally dependent. Interestingly, from the point of view of our analysis, both local dependence and asynchrony are solved via the same technical instrument, which we call <em>conflict sets</em>. More formally, each new test has a conflict set,  which consists of all previously started tests whose outcome is not known (e.g. if there is asynchrony so they are still running), or is known but might have some leverage on the new test (e.g. if there is dependence). We show that computing the FPR of a new test while assuming &#8220;unfavorable&#8221; outcomes of the conflicting tests is the right approach to guaranteeing FDR control (we call this the <em>principle of pessimism</em>).</p>

<p>It is worth pointing out that FDR control under conflict sets has to be more conservative by construction; to account for dependence between tests, as well as the uncertainty about the tests in progress, the FPRs have to be chosen appropriately smaller. That said, our methods are a strict generalization of prior work on online FDR; they interpolate between standard online FDR algorithms, when the conflict sets are empty, and the Bonferroni correction (also known as alpha-spending), when the conflict sets are arbitrarily large. The latter controls the familywise error rate, which is a more stringent error metric than FDR, under any assumption on how tests relate. This interpolation has introduced the possibility of a tradeoff between the consideration of overall rate of discovery per unit of real time, and consideration of the complexity of careful coordination required to minimize dependence and asynchrony.</p>

<h1>Summary</h1>

<p>The replicability of hypothesis tests is largely in crisis, as the scale of modern applications has long outstripped classical testing methodology which is still in use. Moreover, prior efforts toward remedying this problem have neglected the fact that testing is massively asynchronous, and hence the existing solutions for boosting reproducibility have not been suitable for many common large-scale testing schemes. Motivated by this observation, we developed methods that control the false discovery rate in complex asynchronous scenarios, allowing statisticians to perform hypothesis tests with a small fraction of false discoveries, and with minimal explicit coordination between tests.</p>

<h2>References</h2>

<p>[1] Zrnic, T., Ramdas, A., &#38; Jordan, M. I. (2018). <a href="https://arxiv.org/abs/1812.05068" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Asynchronous Online Testing
of Multiple Hypotheses</a>. arXiv preprint arXiv:1812.05068.</p>

<hr><div>
  <ol><li>
      <p>Valid p-values can also be stochastically larger than uniform, which is a more general condition. For simplicity, we take them to be uniform in this text; the &#8220;punchline&#8221; remains the same either way.&#160;<a href="http://bair.berkeley.edu/blog/#fnref:pvalue" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">&#8617;</a></p>
    </li>
  </ol></div>]]></description>
										<content:encoded><![CDATA[<p><strong>By Tijana Zrnic</strong>  </p>
<blockquote><p> “Scientific research has changed the world. Now it needs to change itself. &#8211; The Economist, 2013 </p></blockquote>
<p>There has been a growing concern about the validity of scientific findings. A multitude of journals, papers and reports have recognized the ever smaller number of replicable scientific studies. In 2016, one of the giants of scientific publishing, Nature, surveyed about 1,500 researchers across many different disciplines, asking for their stand on the status of reproducibility in their area of research. One of the many takeaways to the worrisome results of this survey is the following: 90% of the respondents agreed that there is a reproducibility crisis, and the overall top answer to boosting reproducibility was “better understanding of statistics”. Indeed, many factors contributing to the explosion of irreproducible research stem from the neglect of the fact that statistics is no longer as static as it was in the first half of the 20th century, when statistical hypothesis testing came into prominence as a theoretically rigorous proposal for making valid discoveries with high confidence.</p>
<p>  <span id="more-115201"></span>  </p>
<p>When science first saw the rise of statistical testing, the basic idea was the following: you put forward competing hypotheses about the world, then you collect some data, and finally you use these data to validate your hypotheses. Typically, one was in a situation where they could iterate this three-step process only a few times; data was scarce, and the necessary computations were lengthy. Remember, this is early to mid-20th century we are talking about.</p>
<p>This forerunner of today’s scientific investigations would hardly recognize its own field in 2019. Nowadays, testing is much more dynamic and is performed at a scale larger than ever before. Even within a single institution,  thousands of hypotheses are tested in a short time interval, older test results inspire future potential analyses, and scientific exploration oftentimes becomes a never-ending stream of individual hypothesis tests. What enabled this explosion of exploratory research is the high-throughput technologies and large amounts of data that we started seeing only recently, at least relative to the era of statistical thinking.</p>
<p>That said, as in any discipline with well-established and successful foundations, it is difficult to move away from classical paradigms in testing. Much of today’s large-scale investigations still uses tools and techniques which, although powerful and supported by beautiful theory, do not take into account that each test might be just a little piece of a much bigger puzzle of exploratory research. Many disciplines have yet to acquire novel methodology for testing, one that promotes valid inferences <em>at scale</em> and thus limits grandiose publications comprised of irreplicable mirages.</p>
<p>Let us analyze why classical hypothesis testing might lead to many spurious conclusions when the number of tests is large. We do so by elaborating the three main steps of a test: “<em>hypothesize</em>”, “<em>collect data</em>” and “<em>validate</em>”.</p>
<p>In the “hypothesize” step, a well-defined <em>null hypothesis</em> is formulated. For example, this could be “jelly beans do not cause acne”; we will use this as our running example. Notice that the null hypothesis, or simply null, is the opposite of what would be considered a discovery. In short, the null is status quo. Also at the beginning of a test, a false positive rate (FPR) is chosen. This is the maximal allowed probability of making a false discovery, typically chosen around 0.05. In the context of our running example, this means the following: if the null is true, i.e. if jelly beans do <strong>not</strong> cause acne, we will only have a 5% chance of proclaiming causation between jelly beans and acne.</p>
<p>In a frequentist manner, we assume that there is deterministic ground truth about the null hypothesis. That is, it is either true or not. We will refer to the null hypotheses that are true as <em>true nulls</em>, and to those that are false as <em>non-nulls</em>. In our example, if jelly beans do not cause acne, the null hypothesis is a true null. If it is a non-null, however, we would ideally like to proclaim a discovery.</p>
<p>The second step is calculating a <em>p-value</em> based on collected data. This protagonist of many controversies around statistical testing is the probability of seeing the collected data, or something even more extreme, <strong>if the null is true</strong>. In our example, this is the probability of having some observed parameter of skin condition, or something “even more unusual”, if jelly beans indeed do not cause acne. To illustrate this point, consider the plot below. Let the bell curve be the distribution of the skin parameter if jelly beans do not cause acne. Then, the p-value is the red shaded area under this curve, which is everything “right of” the observed data point. The smaller the p-value, the more unlikely it is that the observation can be explained purely by chance.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/false-discoveries/pvalue.jpg" width="600" /> <br /> <i> <!-- [source: https://ottawacitizen.com/technology/science/science-worlds-p-value-controversy-little-number-big-problem] --> <a href="https://ottawacitizen.com/technology/science/science-worlds-p-value-controversy-little-number-big-problem" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">[Source]</a> </i> </p>
<p>The last step is validation. If the calculated p-value is smaller than the FPR, the null hypothesis is <em>rejected</em>, and a discovery is proclaimed. In our running example, if the red shaded area is less than 0.05, we say that jelly beans cause acne.</p>
<p>Finally, let us lift the lid on why there are so many false discoveries in large-scale testing. By construction, valid p-values are uniformly distributed on $[0,1]$<sup id="fnref:pvalue"><a href="http://bair.berkeley.edu/blog/2019/02/15/false-discoveries/#fn:pvalue" class="footnote" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">1</a></sup>, <strong>if the null is true</strong>. This means that, even if jelly beans do not really cause acne, there is still 0.05 probability that a discovery is falsely proclaimed. Therefore, if testing N hypotheses that are truly null and hence should <em>not</em> be discovered, one is almost certain to proclaim some of them as discoveries if $N$ is large.  For example, if all tests are independent, around 5% of $N$ will be discovered.  Already after 20 tests of true nulls, even if they are completely arbitrary, one is expected to make a false discovery!</p>
<p>And this is how science goes wrong.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/false-discoveries/significant.png" /> <br /> <i> <a href="https://xkcd.com/882/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">[Source]</a> </i> </p>
<p>To recap, around 5% of the tested <em>true null hypotheses</em> unfortunately have to be discovered either way, simply by laws of probability. This wouldn’t really be an issue if most of the tested hypotheses were legitimate potential discoveries, i.e. non-nulls. Then, 5% of a small-ish number of true nulls would be negligible. Typically, however, this is not the case. We test loads of crazy, out-there hypotheses, which would attract a lot of attention if confirmed, and we do so simply because we can. In many areas, both observations and computational resources are abundant, so there is little incentive to stay on the “safe side”.</p>
<p>So, how can one make scientific discoveries without the fear of reporting too many false ones?</p>
<h1 id="controlling-the-false-discovery-rate">Controlling the False Discovery Rate</h1>
<p>The recognition that a large number of tests leads to almost sure false discoveries has led to various formalisms for controlling their rate of appearance. One powerful proposal, which has become a de facto standard for false discovery control in multiple testing, is called the <strong>false discovery rate</strong> (FDR), defined as:</p>
<p>  <script type="math/tex; mode=display">\text{FDR} = \mathbf{E}\left[\frac{\# \text{ false discoveries}}{\# \text{ discoveries} \vee 1}\right].</script>  </p>
<p>Controlling FDR with no additional goal is an easy task; namely, making no discoveries trivially gives FDR = 0. The implict goal behind the vast literature on FDR is <em>discovering as many non-nulls as possible, while keeping FDR controlled under a pre-specified level $\alpha$</em>. We collectively refer to all methods with this goal as FDR methods.</p>
<p>Initially, FDR methods were offline procedures. This means that they required collecting a whole batch of p-values before deciding which tests to proclaim as discoveries. The most notable example of this class is the successful Benjamini-Hochberg procedure, which has for a long time been the default of FDR methods.</p>
<p>However, the scale and scope of modern testing have begun to outstrip this well-recognized methodology. It is far from convenient to wait for all the p-values one wants to test, especially at institutions where testing is a never-ending process. To be more precise, we typically want to make decisions during and between our tests, in particular because this allows us to shape future analyses based on outcomes of past tests. This inspired a new line of work on FDR control, in which decisions are made <em>online</em>.</p>
<p>In online FDR control, p-values arrive one at a time, and the decision of whether or not to make a discovery is made as soon as a p-value is observed. Importantly, online FDR algorithms have enabled controlling FDR over a lifetime; even if the number of sequential tests tends to infinity, one would still have a guarantee that most of the proclaimed discoveries are indeed non-nulls.</p>
<p>The basic principle of online FDR control is to track and control a dynamic quantity called <em>wealth</em>. The wealth represents the current error budget, and is a result of all previously performed tests. In particular, if a test results in a discovery, the wealth increases, while if a discovery is not made, the wealth decreases; note that this update is completely independent of whether the test is truly null or not. When a new test starts, its FPR is chosen based on the available wealth; the bigger the wealth, the bigger the FPR, and consequently the better the chance for a discovery. In fact, this idea has a perfect analogy with testing in a broader social context. To make scientific discoveries, you are awarded an initial grant (corresponding to the target FDR level $\alpha$). This initial funding decreases with every new experiment, and, if you happen to make a scientific discovery, you are again awarded some “wealth”, which you can use toward the budget for subsequent tests. This is essentially the real-world translation of the mathematical expressions guiding online FDR algorithms.</p>
<h1 id="asynchronous-control-of-false-discoveries">Asynchronous Control of False Discoveries</h1>
<p>Although online FDR control has broadened the domain of applications where false discoveries can be controlled, it has failed to account for several important aspects of modern testing.</p>
<p>The main observation is that large-scale testing is not only sequential, but “doubly sequential”. Tests are run in a sequential fashion, but also each test internally is comprised of a sequence of atomic executions, which typically finish at an unpredictable time. This fact makes practitioners run multiple tests that overlap in time in order to gain time efficiency, allowing tests to start and finish at random times.</p>
<p>For example, in clinical trials, it is common to test several different treatment variants against a common control. These trials are often called “perpetual’’, as multiple treatments are tested in parallel, and new treatments enter the testing platform at random times in an online manner. Similarly, A/B testing in industry is typically distributed across many individuals and research teams, and across time, with companies running hundreds of tests per day. This large volume of tests, as well as their complex distribution across many analysts, inevitably causes asynchrony in testing.</p>
<p>This circumstance is a problem for standard online FDR methodology. Namely, all existing online FDR algorithms assume tests are run synchronously, with <strong>no overlap in time</strong>; in other words, in order to determine a false positive rate for an upcoming test, online FDR methods need to know the outcomes of all previously started tests. The figure below depicts the difference between synchronous and asynchronous online testing.</p>
<p style="text-align:left;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/false-discoveries/retreatsync.png" />  </p>
<p style="text-align:left;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/false-discoveries/retreatasync.png" width="600" /> <br /> <i> For each time step $t$, $W_t$, $P_t$ and $\alpha_t$ are respectively the available wealth at the beginning of the $(t+1)$-th test, the p-value resulting from the $t$-th test, and the FPR of the $t$-th test. </i> </p>
<p>Furthermore, the asynchronous nature of modern testing introduces patterns of dependence between p-values that do not conform to common assumptions. Prior work on online FDR either assumes perfect independence between p-values (overly optimistic), or arbitrary dependence between all tested p-values in the sequence (overly pessimistic). As data are commonly shared across different tests, the first assumption in clearly difficult to satisfy. In clinical trials, having a common control arm induces dependence; in A/B testing, many tests reuse data from the same shared pool, again causing dependence. On the other end, it is not natural to assume that dependence spills over the entire p-value sequence; older data and test outcomes with time become “stale,” and no longer have direct influence on newly created tests. Modern testing calls for an intermediate notion of dependence, called local dependence, one that assumes p-values that are far enough in the sequence are independent, while any two close enough are likely to depend on each other.</p>
<p>In a recent manuscript [1], we developed FDR methods that confront both of these difficulties of large-scale testing. Our methods control FDR in sequential settings that are arbitrarily asynchronous, and/or yield p-values that are locally dependent. Interestingly, from the point of view of our analysis, both local dependence and asynchrony are solved via the same technical instrument, which we call <em>conflict sets</em>. More formally, each new test has a conflict set,  which consists of all previously started tests whose outcome is not known (e.g. if there is asynchrony so they are still running), or is known but might have some leverage on the new test (e.g. if there is dependence). We show that computing the FPR of a new test while assuming “unfavorable” outcomes of the conflicting tests is the right approach to guaranteeing FDR control (we call this the <em>principle of pessimism</em>).</p>
<p>It is worth pointing out that FDR control under conflict sets has to be more conservative by construction; to account for dependence between tests, as well as the uncertainty about the tests in progress, the FPRs have to be chosen appropriately smaller. That said, our methods are a strict generalization of prior work on online FDR; they interpolate between standard online FDR algorithms, when the conflict sets are empty, and the Bonferroni correction (also known as alpha-spending), when the conflict sets are arbitrarily large. The latter controls the familywise error rate, which is a more stringent error metric than FDR, under any assumption on how tests relate. This interpolation has introduced the possibility of a tradeoff between the consideration of overall rate of discovery per unit of real time, and consideration of the complexity of careful coordination required to minimize dependence and asynchrony.</p>
<h1 id="summary">Summary</h1>
<p>The replicability of hypothesis tests is largely in crisis, as the scale of modern applications has long outstripped classical testing methodology which is still in use. Moreover, prior efforts toward remedying this problem have neglected the fact that testing is massively asynchronous, and hence the existing solutions for boosting reproducibility have not been suitable for many common large-scale testing schemes. Motivated by this observation, we developed methods that control the false discovery rate in complex asynchronous scenarios, allowing statisticians to perform hypothesis tests with a small fraction of false discoveries, and with minimal explicit coordination between tests.</p>
<p>This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
<h2 id="references">References</h2>
<p>[1] Zrnic, T., Ramdas, A., &amp; Jordan, M. I. (2018). <a href="https://arxiv.org/abs/1812.05068" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Asynchronous Online Testing of Multiple Hypotheses</a>. arXiv preprint arXiv:1812.05068.</p>
<hr />
<div class="footnotes">
<ol>
<li id="fn:pvalue">
<p>Valid p-values can also be stochastically larger than uniform, which is a more general condition. For simplicity, we take them to be uniform in this text; the “punchline” remains the same either way.&nbsp;<a href="http://bair.berkeley.edu/blog/2019/02/15/false-discoveries/#fnref:pvalue" class="reversefootnote" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">&#8617;</a></p>
</li>
</ol></div>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Learning preferences by looking at the world</title>
		<link>https://robohub.org/learning-preferences-by-looking-at-the-world/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Tue, 12 Feb 2019 23:04:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/learning-preferences-by-looking-at-the-world/</guid>

					<description><![CDATA[<p>It would be great if we could all have household robots do our chores for us.
Chores are tasks that we want done to make our houses cater more to our
preferences; they are a way in which we want our house to be <em>different</em> from
the way it currently is. However, most &#8220;different&#8221; states are not very
desirable:</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/preferences/different.png"><br></p>
&#60;!--
<i><a href="http://webcomicname.com/post/152958755984" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Source</a></i>
--&#62;

<p>Surely our robot wouldn&#8217;t be so dumb as to go around breaking stuff when we ask
it to clean our house? Unfortunately, <strong>AI systems trained with <a href="https://en.wikipedia.org/wiki/Reinforcement_learning" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reinforcement
learning</a> only optimize features specified in the reward function</strong> and are
indifferent to anything we might&#8217;ve inadvertently left out. Generally, it is
easy to get the reward wrong by forgetting to include preferences for things
that should stay the same, since we are so used to having these preferences
satisfied, and there are <em>so many of them</em>. Consider the room below, and imagine
that we want a robot waiter that serves people at the dining table efficiently.
We might implement this using a reward function that provides 1 reward whenever
the robot serves a dish, and use discounting so that the robot is incentivized
to be efficient. What could go wrong with such a reward function? How would we
need to modify the reward function to take this into account? Take a minute to
think about it.</p>

<!--more-->

<p>
    <img src="http://bair.berkeley.edu/static/blog/preferences/fancy-room.png" width="600"><br></p>

<p>Here&#8217;s an incomplete list we came up with:</p>

<ul><li>The robot might track dirt and oil onto the pristine furniture while serving
food, even if it could clean itself up, because there&#8217;s no reason to clean but
there is a reason to hurry.</li>
  <li>In its hurry to deliver dishes, the robot might knock over the cabinet of wine
bottles, or slide plates to people and knock over the glasses.</li>
  <li>In case of an emergency, such as the electricity going out, we don&#8217;t want the
robot to keep trying to serve dishes &#8211; it should at least be out of the way,
if not trying to help us.</li>
  <li>The robot may serve empty or incomplete dishes, dishes that no one at the
table wants, or even split apart dishes into smaller dishes so there are more
of them.</li>
</ul><p>Note that we&#8217;re not talking about problems with robustness and distributional
shift: while those problems are worth tackling, the point is that <em>even if</em> we
achieve robustness, the simple reward function still incentivizes the above
unwanted behaviors.</p>

<p>It&#8217;s common to hear the informal solution that the robot should try to minimize
its impact on the environment, while still accomplishing the task. This could
potentially allow us to avoid the first three problems above, though the last
one still remains as an example of <a href="https://vkrakovna.wordpress.com/2018/04/02/specification-gaming-examples-in-ai/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">specification gaming</a>. This idea leads to
<a href="https://vkrakovna.wordpress.com/2018/06/05/measuring-and-avoiding-side-effects-using-relative-reachability/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">impact</a> <a href="https://www.alignmentforum.org/posts/yEa7kwoMpsBgaBCgb/towards-a-new-impact-measure" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">measures</a> that attempt to quantify the &#8220;impact&#8221; that an agent
has, typically by looking at the difference between what actually happened and
what would have happened had the robot done nothing. However, this also
penalizes things we want the robot to do. For example, if we ask our robot to
get us coffee, it might buy coffee rather than making coffee itself, because
that would have &#8220;impact&#8221; on the water, the coffee maker, etc. Ultimately, we&#8217;d
like to only prevent <em>negative</em> impacts, which means that we need our AI to have
a better idea of what the <em>right</em> reward function is.</p>

<p>Our key insight is that while it might be hard for humans to make their
preferences explicit, some preferences are implicit in the way the world looks:
<strong>the world state is a result of humans having acted to optimize their
preferences</strong>. This explains why we often want the robot to by default &#8220;do
nothing&#8221; &#8211; if we have already optimized the world state for our preferences,
then most ways of changing it will be bad, and so doing nothing will often
(though not always) be one of the better options available to the robot.</p>

<p>Since the world state is a result of optimization for human preferences, we
should be able to use that state to infer what humans care about. For example,
we surely don&#8217;t want dirty floors in our pristine room; otherwise we would have
done that ourselves. We also can&#8217;t be indifferent to dirty floors, because then
at some point we would have walked around the room with dirty shoes and gotten a
dirty floor. The only explanation is that we want the floor to be clean.</p>

<h1>A simple setting</h1>

<p>Let&#8217;s see if we can apply this insight in the simplest possible setting:
gridworlds with a small number of states, a small number of actions, a known
dynamics model (i.e. a model of &#8220;how the world works&#8221;), but an incorrect reward
function. This is a simple enough setting that our robot understands all of the
consequences of its actions. Nevertheless, the problem remains: while the robot
understands <em>what</em> will happen, it still cannot distinguish good consequences
from bad ones, since its reward function is incorrect. In these simple
environments, it&#8217;s easy to figure out what the correct reward function is, but
this is infeasible in a real, complex environment.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/preferences/room.png" hspace="30" align="right" width="240">&#60;!-- <br> --&#62;</p>

<p>For example, consider the room to the right, where Alice asks her robot to
navigate to the purple door. If we were to encode this as a reward function that
only rewards the robot while it is at the purple door, the robot would take the
shortest path to the purple door, knocking over and breaking the vase &#8211; since
no one said it shouldn&#8217;t do that. The robot is perfectly aware that its plan
causes it to break the vase, but by default it doesn&#8217;t realize that it
<em>shouldn&#8217;t</em> break the vase.</p>

<p>In this environment, does it help us to realize that Alice was optimizing the
state of the room for her preferences? Well, if Alice didn&#8217;t care about whether
the vase was broken, she would have probably broken it some time in the past. If
she <em>wanted</em> the vase broken, she definitely would have broken it some time in
the past. So the only consistent explanation is that Alice cared about the vase
being intact, as illustrated in the gif below.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/preferences/overview.gif"><br></p>

<p>While this example has the robot infer that it shouldn&#8217;t take the action of
breaking a vase, the robot can also infer goals that it should actively pursue.
For example, if the robot observes a basket of apples near an apple tree, it can
reasonably infer that Alice wants to harvest apples, since the apples didn&#8217;t
walk into the basket themselves &#8211; Alice must have put effort into picking the
apples and placing them in the basket.</p>

<h1>Reward Learning by Simulating the Past</h1>

<p>We formalize this idea by considering an MDP in which our robot observes the
initial state $s_0$ at deployment, and assumes that it is the result of a human
optimizing some unknown reward for $T$ timesteps.</p>

<p>Before we get to our actual algorithm, consider a completely intractable
algorithm that should do well: for each possible reward function, simulate the
trajectories that Alice would take if she had that reward, and see if the
resulting states are compatible with $s_0$. This set of compatible reward
functions give the candidates for Alice&#8217;s reward function. This is the algorithm
that we implicitly use in the gif above.</p>

<p>Intuitively, this works because:</p>

<ul><li>Anything that requires effort on Alice&#8217;s part (e.g. keeping a vase intact)
will not happen for the vast majority of reward functions, and will force the
reward functions to incentivize that behavior (e.g. by rewarding intact
vases).</li>
  <li>Anything that does not require effort on Alice&#8217;s part (e.g. a vase becoming
dusty) will happen for most reward functions, and so the inferred reward
functions need not incentivize that behavior (e.g. there&#8217;s no particular value
on dusty/clean vases).</li>
</ul><p>Another way to think of it is that we can consider all possible past
trajectories that are compatible with $s_0$, infer the reward function that
makes those trajectories most likely, and keep those reward functions as
plausible candidates, weighted by the number of past trajectories they explain.
Such an algorithm should work for similar reasons. Phrased this way, it sounds
like we want to use <a href="https://people.eecs.berkeley.edu/~russell/papers/colt98-uncertainty.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">inverse reinforcement learning</a> to infer rewards for
every possible past trajectory, and aggregate the results. This is still
intractable, but it turns out we can take this insight and turn it into a
tractable algorithm.</p>

<p>We follow <a href="http://www.cs.cmu.edu/~bziebart/publications/maximum-causal-entropy.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Maximum Causal Entropy Inverse Reinforcement Learning</a> (MCEIRL), a
commonly used algorithm for small MDPs. In this framework, we know the action
space and dynamics of the MDP, as well as a set of good features of the state,
and the reward is assumed to be linear in these features. In addition, the human
is modelled as Boltzmann-rational: Alice&#8217;s probability of taking a particular
action from a given state is assumed to be proportional to the exponent of the
state-action value function Q, computed using soft value iteration. Given these
assumptions, we can calculate $p(\tau \mid \theta_A)$, the distribution over the
possible trajectories $\tau = s_{-T} a_{-T} \dots s_{-1} a_{-1} s_0$ under the
assumption that Alice&#8217;s reward was $\theta_A$. MCEIRL then finds the $\theta_A$
that maximizes the probability of a set of trajectories .</p>

<p>Rather than considering all possible trajectories and running MCEIRL on all of
them to maximize each of their probabilities individually, we instead maximize
the probability of the evidence that we see: the single state $s_0$. To get a
distribution over $s_0$, we marginalize out the human&#8217;s behavior prior to the
robot&#8217;s initialization:</p>

<p>We then find a reward $\theta_A$ that maximizes the likelihood above using
gradient ascent, where the gradient is analytically computed using dynamic
programming. We call this algorithm <em>Reward Learning by Simulating the Past
(RLSP)</em> since it infers the unknown human reward from a single state by
considering what must have happened in the past.</p>

<h1>Using the inferred reward</h1>

<p>While RLSP infers a reward that captures the information about human preferences
contained in the initial state, it is not clear how we should <em>use</em> that reward.
This is a challenging problem &#8211; we have two sources of information, the
inferred reward from $s_0$, and the specified reward $\theta_{\text{spec}}$, and
they will conflict. If Alice has a messy room, $\theta_A$ is not going to
incentivize cleanliness, even though $\theta_{\text{spec}}$ might.</p>

<p>Ideally, we would note the scenarios under which the two rewards conflict, and
ask Alice how she would like to proceed. However, in this work, to demonstrate
the algorithm we use the simple heuristic of adding the two rewards, giving us a
final reward $\theta_A + \lambda \theta_{\text{spec}}$, where $\lambda$ is a
hyperparameter that controls the tradeoff between the rewards.</p>

<p>We designed a suite of simple gridworlds to showcase the properties of RLSP. The
top row shows the behavior when optimizing the (incorrect) specified reward,
while the bottom row shows the behavior you get when you take into account the
reward inferred by RLSP. A more thorough description of each environment is
given in the paper. The last environment in particular shows a limitation of our
method. In a room where the vase is far away from Alice&#8217;s most probable
trajectories, the only trajectories that Alice could have taken to break the
vase are all very long and contribute little to the RLSP likelihood. As a
result, observing the intact vase doesn&#8217;t tell the robot much about whether
Alice wanted to actively avoid breaking the vase, since she wouldn&#8217;t have been
likely to break it in any case.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/preferences/all-environments.gif"><br></p>

<h1>What&#8217;s next?</h1>

<p>Now that we have a basic algorithm that can learn the human preferences from one
state, the natural next step is to scale it to realistic environments where the
states cannot be enumerated, the dynamics are not known, and the reward function
is not linear.  This could be done by adapting existing inverse RL algorithms,
similarly to how we adapted Maximum Causal Entropy IRL to the one-state setting.</p>

<p>The unknown dynamics setting, where we don&#8217;t know &#8220;how the world works&#8221;, is
particularly challenging. Our algorithm relies heavily on the assumption that
our robot knows how the world works &#8211; this is what gives it the ability to
simulate what Alice &#8220;must have done&#8221; in the past. We certainly can&#8217;t learn how
the world works just by observing a single state of the world, so we would have
to learn a dynamics model while acting that can then be used to simulate the
past (and these simulations will get better as the model gets better).</p>

<p>Another avenue for future work is to investigate the ways to decompose the
inferred reward into $\theta_{A, \text{task}}$ which says which task Alice is
performing (&#8220;go to the black door&#8221;), and $\theta_{\text{frame}}$, which captures
what Alice prefers to keep unchanged (&#8220;don&#8217;t break the vase&#8221;). Given the
separate $\theta_{\text{frame}}$, the robot could optimize
$\theta_{\text{spec}}+\theta_{\text{frame}}$ and ignore the parts of the reward
function that correspond to the task Alice is trying to perform.</p>

<p>Since $\theta_{\text{frame}}$ is in large part shared across many humans, we
could infer it using models where multiple humans are optimizing their own
unique $\theta_{H,\text{task}}$ but the same $\theta_{\text{frame}}$, or we
could have one human whose task change over time. Another direction would be to
assume a different structure for what Alice prefers to keep unchanged, such as
constraints, and learn them separately.</p>

<p>You can learn more about this research by reading <a href="https://openreview.net/forum?id=rkevMnRqYQ&#038;noteId=r1eINIUbe4" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our paper</a>, or by checking
out our poster at ICLR 2019. The code is available <a href="https://github.com/HumanCompatibleAI/rlsp" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>.</p>]]></description>
										<content:encoded><![CDATA[<p>By Rohin Shah and Dmitrii Krasheninnikov </p>
<p>It would be great if we could all have household robots do our chores for us. Chores are tasks that we want done to make our houses cater more to our preferences; they are a way in which we want our house to be <em>different</em> from the way it currently is. However, most “different” states are not very desirable:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/preferences/different.png" />  </p>
<p> <!-- <i><a href="http://webcomicname.com/post/152958755984" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Source</a></i> -->  </p>
<p>Surely our robot wouldn’t be so dumb as to go around breaking stuff when we ask it to clean our house? Unfortunately, <strong>AI systems trained with <a href="https://en.wikipedia.org/wiki/Reinforcement_learning" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reinforcement learning</a> only optimize features specified in the reward function</strong> and are indifferent to anything we might’ve inadvertently left out. Generally, it is easy to get the reward wrong by forgetting to include preferences for things that should stay the same, since we are so used to having these preferences satisfied, and there are <em>so many of them</em>. Consider the room below, and imagine that we want a robot waiter that serves people at the dining table efficiently. We might implement this using a reward function that provides 1 reward whenever the robot serves a dish, and use discounting so that the robot is incentivized to be efficient. What could go wrong with such a reward function? How would we need to modify the reward function to take this into account? Take a minute to think about it.</p>
<p>  <span id="more-115111"></span><br />
<img decoding="async" src="https://robohub.org/wp-content/uploads/2019/02/fancy-room-1.png" alt="" width="900" height="599" class="aligncenter size-full wp-image-115133" srcset="https://robohub.org/wp-content/uploads/2019/02/fancy-room-1.png 900w, https://robohub.org/wp-content/uploads/2019/02/fancy-room-1-425x283.png 425w, https://robohub.org/wp-content/uploads/2019/02/fancy-room-1-768x511.png 768w" sizes="(max-width: 900px) 100vw, 900px" /></p>
<p>Here’s an incomplete list we came up with:</p>
<ul>
<li>The robot might track dirt and oil onto the pristine furniture while serving food, even if it could clean itself up, because there’s no reason to clean but there is a reason to hurry.</li>
<li>In its hurry to deliver dishes, the robot might knock over the cabinet of wine bottles, or slide plates to people and knock over the glasses.</li>
<li>In case of an emergency, such as the electricity going out, we don’t want the robot to keep trying to serve dishes – it should at least be out of the way, if not trying to help us.</li>
<li>The robot may serve empty or incomplete dishes, dishes that no one at the table wants, or even split apart dishes into smaller dishes so there are more of them.</li>
</ul>
<p>Note that we’re not talking about problems with robustness and distributional shift: while those problems are worth tackling, the point is that <em>even if</em> we achieve robustness, the simple reward function still incentivizes the above unwanted behaviors.</p>
<p>It’s common to hear the informal solution that the robot should try to minimize its impact on the environment, while still accomplishing the task. This could potentially allow us to avoid the first three problems above, though the last one still remains as an example of <a href="https://vkrakovna.wordpress.com/2018/04/02/specification-gaming-examples-in-ai/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">specification gaming</a>. This idea leads to <a href="https://vkrakovna.wordpress.com/2018/06/05/measuring-and-avoiding-side-effects-using-relative-reachability/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">impact</a> <a href="https://www.alignmentforum.org/posts/yEa7kwoMpsBgaBCgb/towards-a-new-impact-measure" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">measures</a> that attempt to quantify the “impact” that an agent has, typically by looking at the difference between what actually happened and what would have happened had the robot done nothing. However, this also penalizes things we want the robot to do. For example, if we ask our robot to get us coffee, it might buy coffee rather than making coffee itself, because that would have “impact” on the water, the coffee maker, etc. Ultimately, we’d like to only prevent <em>negative</em> impacts, which means that we need our AI to have a better idea of what the <em>right</em> reward function is.</p>
<p>Our key insight is that while it might be hard for humans to make their preferences explicit, some preferences are implicit in the way the world looks: <strong>the world state is a result of humans having acted to optimize their preferences</strong>. This explains why we often want the robot to by default “do nothing” – if we have already optimized the world state for our preferences, then most ways of changing it will be bad, and so doing nothing will often (though not always) be one of the better options available to the robot.</p>
<p>Since the world state is a result of optimization for human preferences, we should be able to use that state to infer what humans care about. For example, we surely don’t want dirty floors in our pristine room; otherwise we would have done that ourselves. We also can’t be indifferent to dirty floors, because then at some point we would have walked around the room with dirty shoes and gotten a dirty floor. The only explanation is that we want the floor to be clean.</p>
<h1 id="a-simple-setting">A simple setting</h1>
<p>Let’s see if we can apply this insight in the simplest possible setting: gridworlds with a small number of states, a small number of actions, a known dynamics model (i.e. a model of “how the world works”), but an incorrect reward function. This is a simple enough setting that our robot understands all of the consequences of its actions. Nevertheless, the problem remains: while the robot understands <em>what</em> will happen, it still cannot distinguish good consequences from bad ones, since its reward function is incorrect. In these simple environments, it’s easy to figure out what the correct reward function is, but this is infeasible in a real, complex environment.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/preferences/room.png" hspace="30" align="right" width="240" /> <!-- <br /> --> </p>
<p>For example, consider the room to the right, where Alice asks her robot to navigate to the purple door. If we were to encode this as a reward function that only rewards the robot while it is at the purple door, the robot would take the shortest path to the purple door, knocking over and breaking the vase – since no one said it shouldn’t do that. The robot is perfectly aware that its plan causes it to break the vase, but by default it doesn’t realize that it <em>shouldn’t</em> break the vase.</p>
<p>In this environment, does it help us to realize that Alice was optimizing the state of the room for her preferences? Well, if Alice didn’t care about whether the vase was broken, she would have probably broken it some time in the past. If she <em>wanted</em> the vase broken, she definitely would have broken it some time in the past. So the only consistent explanation is that Alice cared about the vase being intact, as illustrated in the gif below.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/preferences/overview.gif" />  </p>
<p>While this example has the robot infer that it shouldn’t take the action of breaking a vase, the robot can also infer goals that it should actively pursue. For example, if the robot observes a basket of apples near an apple tree, it can reasonably infer that Alice wants to harvest apples, since the apples didn’t walk into the basket themselves – Alice must have put effort into picking the apples and placing them in the basket.</p>
<h1 id="reward-learning-by-simulating-the-past">Reward Learning by Simulating the Past</h1>
<p>We formalize this idea by considering an MDP in which our robot observes the initial state $s_0$ at deployment, and assumes that it is the result of a human optimizing some unknown reward for $T$ timesteps.</p>
<p>Before we get to our actual algorithm, consider a completely intractable algorithm that should do well: for each possible reward function, simulate the trajectories that Alice would take if she had that reward, and see if the resulting states are compatible with $s_0$. This set of compatible reward functions give the candidates for Alice’s reward function. This is the algorithm that we implicitly use in the gif above.</p>
<p>Intuitively, this works because:</p>
<ul>
<li>Anything that requires effort on Alice’s part (e.g. keeping a vase intact) will not happen for the vast majority of reward functions, and will force the reward functions to incentivize that behavior (e.g. by rewarding intact vases).</li>
<li>Anything that does not require effort on Alice’s part (e.g. a vase becoming dusty) will happen for most reward functions, and so the inferred reward functions need not incentivize that behavior (e.g. there’s no particular value on dusty/clean vases).</li>
</ul>
<p>Another way to think of it is that we can consider all possible past trajectories that are compatible with $s_0$, infer the reward function that makes those trajectories most likely, and keep those reward functions as plausible candidates, weighted by the number of past trajectories they explain. Such an algorithm should work for similar reasons. Phrased this way, it sounds like we want to use <a href="https://people.eecs.berkeley.edu/~russell/papers/colt98-uncertainty.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">inverse reinforcement learning</a> to infer rewards for every possible past trajectory, and aggregate the results. This is still intractable, but it turns out we can take this insight and turn it into a tractable algorithm.</p>
<p>We follow <a href="http://www.cs.cmu.edu/~bziebart/publications/maximum-causal-entropy.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Maximum Causal Entropy Inverse Reinforcement Learning</a> (MCEIRL), a commonly used algorithm for small MDPs. In this framework, we know the action space and dynamics of the MDP, as well as a set of good features of the state, and the reward is assumed to be linear in these features. In addition, the human is modelled as Boltzmann-rational: Alice’s probability of taking a particular action from a given state is assumed to be proportional to the exponent of the state-action value function Q, computed using soft value iteration. Given these assumptions, we can calculate $p(\tau \mid \theta_A)$, the distribution over the possible trajectories $\tau = s_{-T} a_{-T} \dots s_{-1} a_{-1} s_0$ under the assumption that Alice’s reward was $\theta_A$. MCEIRL then finds the $\theta_A$ that maximizes the probability of a set of trajectories <script type="math/tex">\{\tau_i\}</script>.</p>
<p>Rather than considering all possible trajectories and running MCEIRL on all of them to maximize each of their probabilities individually, we instead maximize the probability of the evidence that we see: the single state $s_0$. To get a distribution over $s_0$, we marginalize out the human’s behavior prior to the robot’s initialization:</p>
<p>  <script type="math/tex; mode=display">P(s_0 \mid \theta_A) = \sum\limits_{s_{-T}a_{-T} \dots s_{-1}a_{-1}} P(s_{-T} a_{-T} \dots s_{-1} a_{-1} s_0 \mid \theta_A)</script>  </p>
<p>We then find a reward $\theta_A$ that maximizes the likelihood above using gradient ascent, where the gradient is analytically computed using dynamic programming. We call this algorithm <em>Reward Learning by Simulating the Past (RLSP)</em> since it infers the unknown human reward from a single state by considering what must have happened in the past.</p>
<h1 id="using-the-inferred-reward">Using the inferred reward</h1>
<p>While RLSP infers a reward that captures the information about human preferences contained in the initial state, it is not clear how we should <em>use</em> that reward. This is a challenging problem – we have two sources of information, the inferred reward from $s_0$, and the specified reward $\theta_{\text{spec}}$, and they will conflict. If Alice has a messy room, $\theta_A$ is not going to incentivize cleanliness, even though $\theta_{\text{spec}}$ might.</p>
<p>Ideally, we would note the scenarios under which the two rewards conflict, and ask Alice how she would like to proceed. However, in this work, to demonstrate the algorithm we use the simple heuristic of adding the two rewards, giving us a final reward $\theta_A + \lambda \theta_{\text{spec}}$, where $\lambda$ is a hyperparameter that controls the tradeoff between the rewards.</p>
<p>We designed a suite of simple gridworlds to showcase the properties of RLSP. The top row shows the behavior when optimizing the (incorrect) specified reward, while the bottom row shows the behavior you get when you take into account the reward inferred by RLSP. A more thorough description of each environment is given in the paper. The last environment in particular shows a limitation of our method. In a room where the vase is far away from Alice’s most probable trajectories, the only trajectories that Alice could have taken to break the vase are all very long and contribute little to the RLSP likelihood. As a result, observing the intact vase doesn’t tell the robot much about whether Alice wanted to actively avoid breaking the vase, since she wouldn’t have been likely to break it in any case.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/preferences/all-environments.gif" />  </p>
<h1 id="whats-next">What’s next?</h1>
<p>Now that we have a basic algorithm that can learn the human preferences from one state, the natural next step is to scale it to realistic environments where the states cannot be enumerated, the dynamics are not known, and the reward function is not linear.  This could be done by adapting existing inverse RL algorithms, similarly to how we adapted Maximum Causal Entropy IRL to the one-state setting.</p>
<p>The unknown dynamics setting, where we don’t know “how the world works”, is particularly challenging. Our algorithm relies heavily on the assumption that our robot knows how the world works – this is what gives it the ability to simulate what Alice “must have done” in the past. We certainly can’t learn how the world works just by observing a single state of the world, so we would have to learn a dynamics model while acting that can then be used to simulate the past (and these simulations will get better as the model gets better).</p>
<p>Another avenue for future work is to investigate the ways to decompose the inferred reward into $\theta_{A, \text{task}}$ which says which task Alice is performing (“go to the black door”), and $\theta_{\text{frame}}$, which captures what Alice prefers to keep unchanged (“don’t break the vase”). Given the separate $\theta_{\text{frame}}$, the robot could optimize $\theta_{\text{spec}}+\theta_{\text{frame}}$ and ignore the parts of the reward function that correspond to the task Alice is trying to perform.</p>
<p>Since $\theta_{\text{frame}}$ is in large part shared across many humans, we could infer it using models where multiple humans are optimizing their own unique $\theta_{H,\text{task}}$ but the same $\theta_{\text{frame}}$, or we could have one human whose task change over time. Another direction would be to assume a different structure for what Alice prefers to keep unchanged, such as constraints, and learn them separately.</p>
<p>You can learn more about this research by reading <a href="https://openreview.net/forum?id=rkevMnRqYQ&amp;noteId=r1eINIUbe4" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our paper</a>, or by checking out our poster at ICLR 2019. The code is available <a href="https://github.com/HumanCompatibleAI/rlsp" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>.</p>
<p>  This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission. </p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Soft actor critic &#8211; Deep reinforcement learning with real-world robots</title>
		<link>https://robohub.org/soft-actor-critic-deep-reinforcement-learning-with-real-world-robots/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sun, 16 Dec 2018 21:15:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/soft-actor-critic-deep-reinforcement-learning-with-real-world-robots/</guid>

					<description><![CDATA[<p>We are announcing the release of our state-of-the-art off-policy model-free
reinforcement learning algorithm, soft actor-critic (SAC). This algorithm has
been developed jointly at UC Berkeley and Google Brain, and we have been using
it internally for our robotics experiment. Soft actor-critic is, to our
knowledge, one of the most efficient model-free algorithms available today,
making it especially well-suited for real-world robotic learning. In this post,
we will benchmark SAC against state-of-the-art model-free RL algorithms and
showcase a spectrum of real-world robot examples, ranging from manipulation to
locomotion. We also release our implementation of SAC, which is particularly
designed for real-world robotic systems.</p>

<!--more-->

<h1>Desired Features for Deep RL for Real Robots</h1>

<p>What makes an ideal deep RL algorithm for real-world systems? Real-world
experimentation brings additional challenges, such as constant interruptions in
the data stream, requirement for a low-latency inference and smooth exploration
to avoid mechanical wear and tear on the robot, which set additional requirement
for both the algorithm and also the implementation of the algorithm.</p>

<p>Regarding the algorithm, several properties are desirable:</p>

<ul><li><strong>Sample Efficiency</strong>. Learning skills in the real world can take a
substantial amount of time. Prototyping a new task takes several trials, and
the total time required to learn a new skill quickly adds up. Thus good sample
complexity is the first prerequisite for successful skill acquisition.</li>
  <li><strong>No Sensitive Hyperparameters</strong>. In the real world, we want to avoid
parameter tuning for the obvious reason. Maximum entropy RL provides a robust
framework that minimizes the need for hyperparameter tuning.</li>
  <li><strong>Off-Policy Learning</strong>. An algorithm is off-policy if we can reuse data collected
for another task. In a typical scenario, we need to adjust parameters and
shape the reward function when prototyping a new task, and use of an
off-policy algorithm allows reusing the already collected data.</li>
</ul><p>Soft actor-critic (SAC), described below, is an off-policy model-free deep RL 
algorithm that is well aligned with these requirements. In particular, we show 
that it is sample efficient enough to solve real-world robot tasks in only a 
handful of hours, robust to hyperparameters and works on a variety of simulated 
environments with a single set of hyperparameters.</p>

<p>In addition to the desired algorithmic properties, experimentation in the
real-world sets additional requirements for the implementation. Our release
supports many of these features that we have found crucial when learning with
real robots, perhaps the most importantly:</p>

<ul><li><strong>Asynchronous Sampling</strong>. Inference needs to be fast to minimize delay in the
control loop, and we typically want to keep training during the environment
resets too. Therefore, data sampling and training should run in independent
threads or processes.</li>
  <li><strong>Stop / Resume Training</strong>. When working with real hardware, whatever can go
wrong, will go wrong. We should expect constant interruptions in the data
stream.</li>
  <li><strong>Action smoothing</strong>. Typical Gaussian exploration makes the actuators jitter
at high frequency, potentially damaging the hardware. Thus temporally
correlating the exploration is important.</li>
</ul><h1>Soft Actor-Critic</h1>

<p>Soft actor-critic is based on the maximum entropy reinforcement learning
framework, which considers the entropy augmented objective</p>

<p>where $\mathbf{s}_t$ and $\mathbf{a}_t$ are the state and the action, and the
expectation is taken over the policy and the true dynamics of the system. In
other words, the optimal policy not only maximizes the expected return (first
summand) but also the expected entropy of itself (second summand). The trade-off
between the two is controlled by the non-negative temperature parameter
$\alpha$, and we can always recover the conventional, maximum expected return
objective by setting $\alpha=0$. In a <a href="https://drive.google.com/open?id=1J8gZXJN0RqH-TkTh4UEikYSy8AqPTy9x" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">technical report</a>, we show that we can 
view this objective as an entropy constrained maximization of the expected 
return, and learn the temperature parameter automatically instead of treating 
it as a hyperparameter.</p>

<p>This objective can be interpreted in several ways. We can view the entropy term
as an uninformative (uniform) prior over the policy, but we can also view it as
a regularizer or as an attempt to trade off between exploration (maximize
entropy) and exploitation (maximize return). In our <a href="https://bair.berkeley.edu/blog/2017/10/06/soft-q-learning/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">previous post</a>, we gave
a broader overview and proposed applications that are unique to maximum entropy
RL, and a probabilistic view of the objective is discussed in a <a href="https://arxiv.org/abs/1805.00909" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recent
tutorial</a>.  Soft actor-critic maximizes this objective by parameterizing a
Gaussian policy and a Q-function with a neural network, and optimizing them
using approximate dynamic programming. We defer further details of soft
actor-critic to the <a href="https://drive.google.com/open?id=1J8gZXJN0RqH-TkTh4UEikYSy8AqPTy9x" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">technical report</a>. In this post, we will view the objective as
a grounded way to derive better reinforcement learning algorithms that perform
consistently and are sample efficient enough to be applicable to real-world
robotic applications, and&#8212;perhaps surprisingly&#8212;can yield state-of-the-art
performance under the conventional, maximum expected return objective (without
entropy regularization) in simulated benchmarks.</p>

<h1>Simulated Benchmarks</h1>

<p>Before we jump into real-world experiments, we compare SAC on standard benchmark
tasks to other popular deep RL algorithms, deep deterministic policy gradient
(DDPG), twin delayed deep deterministic policy gradient (TD3), and proximal
policy optimization (PPO). The figures below compare the algorithms on three
challenging locomotion tasks, HalfCheetah, Ant, and Humanoid, from OpenAI Gym.
The solid lines depict the total average return and the shadings correspond to
the best and the worst trial over five random seeds. Indeed, soft actor-critic,
which is shown in blue, achieves the best performance, and&#8212;what&#8217;s even more
important for real-world applications&#8212;it performs well also in the worst case.
We have included more benchmark results in the <a href="https://drive.google.com/open?id=1J8gZXJN0RqH-TkTh4UEikYSy8AqPTy9x" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">technical report</a>.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/sac/ant.png" height="190"><img src="http://bair.berkeley.edu/static/blog/sac/cheetah.png" height="190"><img src="http://bair.berkeley.edu/static/blog/sac/humanoid.png" height="190"><br></p>

<h1>Deep RL in the Real World</h1>

<p>We tested soft actor-critic in the real world by solving three tasks from 
scratch without relying on simulation or demonstrations. 
Our first real-world task involves the Minitaur robot, a small-scale quadruped
with eight direct-drive actuators. The action space consists of the swing angle
and the extension of each leg, which are then mapped to desired motor positions
and tracked with a PD controller. The observations include the motor angles as
well as roll and pitch angles and angular velocities of the base. This learning
task presents substantial challenges for real-world reinforcement learning. The
robot is underactuated, and must therefore delicately balance contact forces on
the legs to make forward progress. An untrained policy can lose balance and
fall, and too many falls will eventually damage the robot, making
sample-efficient learning essentially. The video below illustrates the learned
skill. Although we trained our policy only on flat terrain, we then tested it on
varied terrains and obstacles. Because soft actor-critic learns robust policies,
due to entropy maximization at training time, the policy can readily generalize
to these perturbations without any additional learning.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/sac/Minitaur-evaluation-all.gif" height="180" width="320"><br><i>
The Minitaur robot (Google Brain, Tuomas Haarnoja, Sehoon Ha, Jie Tan, and
Sergey Levine).
</i>
</p>

<p>Our second real-world robotic task involves training a 3-finger dexterous
robotic hand to manipulate an object. The hand is based on the Dynamixel Claw
hand, discussed in <a href="https://bair.berkeley.edu/blog/2018/08/31/dexterous-manip/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">another post</a>. This hand has 9 DoFs, each controlled by a
Dynamixel servo-motor. The policy controls the hand by sending target joint
angle positions for the on-board PID controller. The manipulation task requires
the hand to rotate a ``valve&#8217;&#8216;-like object as shown in the animation below. In
order to perceive the valve, the robot must use raw RGB images shown in the
inset at the bottom right. The robot must rotate the valve so that the colored 
peg faces the right (see video below). The initial position of the valve is reset
uniformly at random for each episode, forcing the policy to learn to use the raw
RGB images to perceive the current valve orientation. A small motor is attached
to the valve to automate resets and to provide the ground truth position for the
determination of the reward function. The position of this motor is not provided
to the policy. This task is exceptionally challenging due to both the perception
challenges and the need to control a hand with 9 degrees of freedom.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/sac/dclaw_vision_rollouts_combined.gif" height="240" width="320"><br><i>
Rotating a valve with a dexterous hand, learned directly from raw pixels  
(UC Berkeley, Kristian Hartikainen, Vikash Kumar, Henry Zhu, Abhishek Gupta, 
Tuomas Haarnoja, and Sergey Levine).
</i>
</p>

<p>In the final task, we trained a 7-DoF Sawyer robot to stack Lego blocks. The
policy receives the joint positions and velocities, as well as end-effector
force as an input and outputs torque commands to each of the seven joints. The
biggest challenge is to accurately align the studs before exerting a
downward force to overcome the friction between them.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/sac/sawyer.gif" height="240" width="320"><br><i>
Stacking Legos with Sawyer (UC Berkeley, Aurick Zhou, Tuomas Haarnoja, and
Sergey Levine).
</i>
</p>

<p>Soft actor-critic solves all of these tasks quickly: the Minitaur
locomotion and the block-stacking tasks both take 2 hours, and the valve-turning
task from image observations takes 20 hours. We also learned a policy for the
valve-turning task without images by providing the actual valve position as an
observation to the policy. Soft actor-critic can learn this easier version of
the valve task in 3 hours. For comparison, <a href="https://arxiv.org/abs/1810.06045" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">prior work</a> has used PPO to learn
the same task without images in 7.4 hours.</p>

<h1>Conclusion</h1>

<p>Soft actor-critic is a step towards feasible deep RL with real-world robots.
Work still needs to be done to scale up these methods to more challenging tasks,
but we believe we are getting closer to the critical point where deep RL can
become a practical solution for robotic tasks. Meanwhile, you can connect your
robot to our toolbox and get learning started!</p>

<h2>Acknowledgements</h2>

<p>We would like to thank the amazing teams at Google Brain and UC
Berkeley&#8212;specifically Pieter Abbeel, Abhishek Gupta, Sehoon Ha, Vikash Kumar,
Sergey Levine, Jie Tan, George Tucker, Vincent Vanhoucke, Henry Zhu&#8212;who
contributed to the development of the algorithm, spent long days running
experiments, and provided the support and resources that made the project
possible.</p>

<p>Links:</p>
<ul><li><a href="https://sites.google.com/view/sac-and-applications" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project website</a></li>
  <li><a href="https://drive.google.com/open?id=1J8gZXJN0RqH-TkTh4UEikYSy8AqPTy9x" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Technical description of SAC</a></li>
  <li><a href="https://github.com/rail-berkeley/softlearning" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">softlearning</a> (our robot learning toolbox, including a SAC implementation in Tensorflow)</li>
  <li><a href="https://github.com/vitchyr/rlkit" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">rlkit</a> (another SAC implementation from UC Berkeley in PyTorch)</li>
</ul>]]></description>
										<content:encoded><![CDATA[<p><strong>By Tuomas Haarnoja, Vitchyr Pong, Kristian Hartikainen, Aurick Zhou, Murtaza Dalal, and Sergey Levine</strong></p>
<p>We are announcing the release of our state-of-the-art off-policy model-free reinforcement learning algorithm, soft actor-critic (SAC). This algorithm has been developed jointly at UC Berkeley and Google Brain, and we have been using it internally for our robotics experiment. Soft actor-critic is, to our knowledge, one of the most efficient model-free algorithms available today, making it especially well-suited for real-world robotic learning. In this post, we will benchmark SAC against state-of-the-art model-free RL algorithms and showcase a spectrum of real-world robot examples, ranging from manipulation to locomotion. We also release our implementation of SAC, which is particularly designed for real-world robotic systems.</p>
<p>  <span id="more-114048"></span>  </p>
<h1 id="desired-features-for-deep-rl-for-real-robots">Desired Features for Deep RL for Real Robots</h1>
<p>What makes an ideal deep RL algorithm for real-world systems? Real-world experimentation brings additional challenges, such as constant interruptions in the data stream, requirement for a low-latency inference and smooth exploration to avoid mechanical wear and tear on the robot, which set additional requirement for both the algorithm and also the implementation of the algorithm.</p>
<p>Regarding the algorithm, several properties are desirable:</p>
<ul>
<li><strong>Sample Efficiency</strong>. Learning skills in the real world can take a substantial amount of time. Prototyping a new task takes several trials, and the total time required to learn a new skill quickly adds up. Thus good sample complexity is the first prerequisite for successful skill acquisition.</li>
<li><strong>No Sensitive Hyperparameters</strong>. In the real world, we want to avoid parameter tuning for the obvious reason. Maximum entropy RL provides a robust framework that minimizes the need for hyperparameter tuning.</li>
<li><strong>Off-Policy Learning</strong>. An algorithm is off-policy if we can reuse data collected for another task. In a typical scenario, we need to adjust parameters and shape the reward function when prototyping a new task, and use of an off-policy algorithm allows reusing the already collected data.</li>
</ul>
<p>Soft actor-critic (SAC), described below, is an off-policy model-free deep RL  algorithm that is well aligned with these requirements. In particular, we show  that it is sample efficient enough to solve real-world robot tasks in only a  handful of hours, robust to hyperparameters and works on a variety of simulated  environments with a single set of hyperparameters.</p>
<p>In addition to the desired algorithmic properties, experimentation in the real-world sets additional requirements for the implementation. Our release supports many of these features that we have found crucial when learning with real robots, perhaps the most importantly:</p>
<ul>
<li><strong>Asynchronous Sampling</strong>. Inference needs to be fast to minimize delay in the control loop, and we typically want to keep training during the environment resets too. Therefore, data sampling and training should run in independent threads or processes.</li>
<li><strong>Stop / Resume Training</strong>. When working with real hardware, whatever can go wrong, will go wrong. We should expect constant interruptions in the data stream.</li>
<li><strong>Action smoothing</strong>. Typical Gaussian exploration makes the actuators jitter at high frequency, potentially damaging the hardware. Thus temporally correlating the exploration is important.</li>
</ul>
<h1 id="soft-actor-critic">Soft Actor-Critic</h1>
<p>Soft actor-critic is based on the maximum entropy reinforcement learning framework, which considers the entropy augmented objective</p>
<p>  <script type="math/tex; mode=display">J(\pi) = \mathbb{E}_{\pi} \left[ {\sum_{t} r(\mathbf{s}_t, \mathbf{a}_t) - \alpha \log (\pi(\mathbf{a}_t|\mathbf{s}_t))} \right],</script>  </p>
<p>where $\mathbf{s}_t$ and $\mathbf{a}_t$ are the state and the action, and the expectation is taken over the policy and the true dynamics of the system. In other words, the optimal policy not only maximizes the expected return (first summand) but also the expected entropy of itself (second summand). The trade-off between the two is controlled by the non-negative temperature parameter $\alpha$, and we can always recover the conventional, maximum expected return objective by setting $\alpha=0$. In a <a href="https://drive.google.com/open?id=1J8gZXJN0RqH-TkTh4UEikYSy8AqPTy9x" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">technical report</a>, we show that we can  view this objective as an entropy constrained maximization of the expected  return, and learn the temperature parameter automatically instead of treating  it as a hyperparameter.</p>
<p>This objective can be interpreted in several ways. We can view the entropy term as an uninformative (uniform) prior over the policy, but we can also view it as a regularizer or as an attempt to trade off between exploration (maximize entropy) and exploitation (maximize return). In our <a href="https://bair.berkeley.edu/blog/2017/10/06/soft-q-learning/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">previous post</a>, we gave a broader overview and proposed applications that are unique to maximum entropy RL, and a probabilistic view of the objective is discussed in a <a href="https://arxiv.org/abs/1805.00909" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recent tutorial</a>.  Soft actor-critic maximizes this objective by parameterizing a Gaussian policy and a Q-function with a neural network, and optimizing them using approximate dynamic programming. We defer further details of soft actor-critic to the <a href="https://drive.google.com/open?id=1J8gZXJN0RqH-TkTh4UEikYSy8AqPTy9x" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">technical report</a>. In this post, we will view the objective as a grounded way to derive better reinforcement learning algorithms that perform consistently and are sample efficient enough to be applicable to real-world robotic applications, and—perhaps surprisingly—can yield state-of-the-art performance under the conventional, maximum expected return objective (without entropy regularization) in simulated benchmarks.</p>
<h1 id="simulated-benchmarks">Simulated Benchmarks</h1>
<p>Before we jump into real-world experiments, we compare SAC on standard benchmark tasks to other popular deep RL algorithms, deep deterministic policy gradient (DDPG), twin delayed deep deterministic policy gradient (TD3), and proximal policy optimization (PPO). The figures below compare the algorithms on three challenging locomotion tasks, HalfCheetah, Ant, and Humanoid, from OpenAI Gym. The solid lines depict the total average return and the shadings correspond to the best and the worst trial over five random seeds. Indeed, soft actor-critic, which is shown in blue, achieves the best performance, and—what’s even more important for real-world applications—it performs well also in the worst case. We have included more benchmark results in the <a href="https://drive.google.com/open?id=1J8gZXJN0RqH-TkTh4UEikYSy8AqPTy9x" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">technical report</a>.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sac/ant.png" height="190" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sac/cheetah.png" height="190" style="margin: 2px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sac/humanoid.png" height="190" style="margin: 2px;" />      </p>
<h1 id="deep-rl-in-the-real-world">Deep RL in the Real World</h1>
<p>We tested soft actor-critic in the real world by solving three tasks from  scratch without relying on simulation or demonstrations.  Our first real-world task involves the Minitaur robot, a small-scale quadruped with eight direct-drive actuators. The action space consists of the swing angle and the extension of each leg, which are then mapped to desired motor positions and tracked with a PD controller. The observations include the motor angles as well as roll and pitch angles and angular velocities of the base. This learning task presents substantial challenges for real-world reinforcement learning. The robot is underactuated, and must therefore delicately balance contact forces on the legs to make forward progress. An untrained policy can lose balance and fall, and too many falls will eventually damage the robot, making sample-efficient learning essentially. The video below illustrates the learned skill. Although we trained our policy only on flat terrain, we then tested it on varied terrains and obstacles. Because soft actor-critic learns robust policies, due to entropy maximization at training time, the policy can readily generalize to these perturbations without any additional learning.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sac/Minitaur-evaluation-all.gif" height="180" width="320" />     <br /> <i> The Minitaur robot (Google Brain, Tuomas Haarnoja, Sehoon Ha, Jie Tan, and Sergey Levine). </i> </p>
<p>Our second real-world robotic task involves training a 3-finger dexterous robotic hand to manipulate an object. The hand is based on the Dynamixel Claw hand, discussed in <a href="https://bair.berkeley.edu/blog/2018/08/31/dexterous-manip/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">another post</a>. This hand has 9 DoFs, each controlled by a Dynamixel servo-motor. The policy controls the hand by sending target joint angle positions for the on-board PID controller. The manipulation task requires the hand to rotate a &#8220;valve’‘-like object as shown in the animation below. In order to perceive the valve, the robot must use raw RGB images shown in the inset at the bottom right. The robot must rotate the valve so that the colored  peg faces the right (see video below). The initial position of the valve is reset uniformly at random for each episode, forcing the policy to learn to use the raw RGB images to perceive the current valve orientation. A small motor is attached to the valve to automate resets and to provide the ground truth position for the determination of the reward function. The position of this motor is not provided to the policy. This task is exceptionally challenging due to both the perception challenges and the need to control a hand with 9 degrees of freedom.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sac/dclaw_vision_rollouts_combined.gif" height="240" width="320" />     <br /> <i> Rotating a valve with a dexterous hand, learned directly from raw pixels   (UC Berkeley, Kristian Hartikainen, Vikash Kumar, Henry Zhu, Abhishek Gupta,  Tuomas Haarnoja, and Sergey Levine). </i> </p>
<p>In the final task, we trained a 7-DoF Sawyer robot to stack Lego blocks. The policy receives the joint positions and velocities, as well as end-effector force as an input and outputs torque commands to each of the seven joints. The biggest challenge is to accurately align the studs before exerting a downward force to overcome the friction between them.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sac/sawyer.gif" height="240" width="320" />     <br /> <i> Stacking Legos with Sawyer (UC Berkeley, Aurick Zhou, Tuomas Haarnoja, and Sergey Levine). </i> </p>
<p>Soft actor-critic solves all of these tasks quickly: the Minitaur locomotion and the block-stacking tasks both take 2 hours, and the valve-turning task from image observations takes 20 hours. We also learned a policy for the valve-turning task without images by providing the actual valve position as an observation to the policy. Soft actor-critic can learn this easier version of the valve task in 3 hours. For comparison, <a href="https://arxiv.org/abs/1810.06045" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">prior work</a> has used PPO to learn the same task without images in 7.4 hours.</p>
<h1 id="conclusion">Conclusion</h1>
<p>Soft actor-critic is a step towards feasible deep RL with real-world robots. Work still needs to be done to scale up these methods to more challenging tasks, but we believe we are getting closer to the critical point where deep RL can become a practical solution for robotic tasks. Meanwhile, you can connect your robot to our toolbox and get learning started!</p>
<h2 id="acknowledgements">Acknowledgements</h2>
<p>We would like to thank the amazing teams at Google Brain and UC Berkeley—specifically Pieter Abbeel, Abhishek Gupta, Sehoon Ha, Vikash Kumar, Sergey Levine, Jie Tan, George Tucker, Vincent Vanhoucke, Henry Zhu—who contributed to the development of the algorithm, spent long days running experiments, and provided the support and resources that made the project possible.</p>
<p> This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission. </p>
<p>Links:</p>
<ul>
<li><a href="https://sites.google.com/view/sac-and-applications" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project website</a></li>
<li><a href="https://drive.google.com/open?id=1J8gZXJN0RqH-TkTh4UEikYSy8AqPTy9x" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Technical description of SAC</a></li>
<li><a href="https://github.com/rail-berkeley/softlearning" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">softlearning</a> (our robot learning toolbox, including a SAC implementation in Tensorflow)</li>
<li><a href="https://github.com/vitchyr/rlkit" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">rlkit</a> (another SAC implementation from UC Berkeley in PyTorch)</li>
</ul>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Visual model-based reinforcement learning as a path towards generalist robots</title>
		<link>https://robohub.org/visual-model-based-reinforcement-learning-as-a-path-towards-generalist-robots/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Tue, 04 Dec 2018 11:14:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/visual-model-based-reinforcement-learning-as-a-path-towards-generalist-robots/</guid>

					<description><![CDATA[With very little explicit supervision and feedback, humans are able to learn a
wide range of motor skills by simply interacting with and observing the world
through their senses. While there has been significant progress towards building
machines that ...]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" src="https://robohub.org/wp-content/uploads/2018/11/Bairfolding.png" alt="" width="900" height="451" class="aligncenter size-full wp-image-113760" srcset="https://robohub.org/wp-content/uploads/2018/11/Bairfolding.png 900w, https://robohub.org/wp-content/uploads/2018/11/Bairfolding-425x213.png 425w, https://robohub.org/wp-content/uploads/2018/11/Bairfolding-768x385.png 768w" sizes="(max-width: 900px) 100vw, 900px" /><br />
<strong>By Chelsea Finn∗, Frederik Ebert∗, Sudeep Dasari, Annie Xie, Alex Lee, and Sergey Levine</strong></p>
<p>With very little explicit supervision and feedback, humans are able to learn a wide range of motor skills by simply interacting with and observing the world through their senses. While there has been significant progress towards building machines that can <a href="https://deepmind.com/research/alphago/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">learn complex skills</a> and <a href="https://www.nytimes.com/2015/05/22/science/robots-that-can-match-human-dexterity.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">learn based on raw sensory information</a> such as image pixels, acquiring large and diverse repertoires of <em>general</em> skills remains an open challenge. Our goal is to build a generalist: a robot that can perform many different tasks, like arranging objects, picking up toys, and folding towels, and can do so with many different objects in the real world without re-learning for each object or task.<span id="more-113710"></span></p>
<p>While these basic motor skills are much simpler and less impressive than <a href="https://arxiv.org/abs/1712.01815" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">mastering Chess</a> or even <a href="https://youtu.be/Z_9yWIe0vCs?t=127" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">using a spatula</a>, we think that being able to achieve such generality with a single model is a fundamental aspect of intelligence.</p>
<p>The key to acquiring generality is <em>diversity</em>. If you deploy a learning algorithm in a narrow, closed-world environment, the agent will recover skills that are successful only in a narrow range of settings. That’s why an algorithm trained to play Breakout will struggle when anything about the images or the game changes. Indeed, the success of image classifiers relies on large, diverse datasets like ImageNet. However, having a robot autonomously learn from large and diverse datasets is quite challenging. While collecting diverse sensory data is relatively straightforward, it is simply not practical for a person to annotate all of the robot’s experiences. It is more scalable to collect completely unlabeled experiences. Then, given only sensory data, akin to what humans have, what can you learn? With raw sensory data there is no notion of progress, reward, or success. Unlike games like Breakout, the real world doesn’t give us a score or extra lives.</p>
<p>We have developed an algorithm that can learn a general-purpose predictive model using unlabeled sensory experiences, and then use this single model to perform a wide range of tasks.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/ezgif-4-e1029b7fe503.gif" height="200" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/ezgif-4-8e9ca4f35be8.gif" height="200" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/ezgif-4-9d83be4e2f4a.gif" height="200" style="margin: 5px;" />     <br />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/ezgif-4-277f7cf520d3.gif" height="200" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/ezgif-4-fe5c96d7f2a8.gif" height="200" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/ezgif-4-f8c4a8d0a89f.gif" height="200" style="margin: 5px;" />     <br /> <i> With a single model, our approach can perform a wide range of tasks, including lifting objects, folding shorts, placing an apple onto a plate, rearranging objects, and covering a fork with a towel. </i> </p>
<p>In this post, we will describe how this works. We will discuss how we can learn based on only raw sensory interaction data (i.e. image pixels, without requiring object detectors or hand-engineered perception components). We will show how we can use what was learned to accomplish many different user-specified tasks. And, we will demonstrate how this approach can control a real robot from raw pixels, performing tasks and interacting with objects that the robot has never seen before.</p>
<p>  <!--more-->  </p>
<h2 id="learning-to-predict-from-unsupervised-interaction">Learning to Predict from Unsupervised Interaction</h2>
<p>We first need a means to collect <em>diverse</em> data. If we train the robot to perform a single skill with a single object instance, i.e. using a particular hammer to hit a particular nail, then it will only learn about that narrow setting; that particular hammer and nail is its entire universe. How can we build robots that learn more general skills? Instead of learning a single task in a narrow environment, we can have robots learn on their own, in diverse environments, akin to <a href="https://youtu.be/8vNxjwt2AqY?t=16" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">a child playing and exploring</a>.</p>
<p>If a robot can collect data on its own and learn from that experience completely autonomously, then it doesn’t require a person to supervise and can hence collect experience and learn about the world at any time of day, even overnight! Further, multiple robots can collect data simultaneously and share their experiences – data collection is <em>scalable</em>, hence making it practical to collect diverse data with many objects and motions. To implement this, we had two robots collect data in parallel by taking random actions with a wide range of objects, both rigid objects like toys and cups, and deformable objects like cloth and towels:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/rigid_obj_experience.gif" height="240" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/pasted-movie-48957.gif" height="240" style="margin: 5px;" />     <br /> <i> Two robots interact with the world, collecting data<br /> autonomously with many objects and many motions. </i> </p>
<p>In the data collection process, we observe what the robot’s sensors measure: the image pixels (vision), the position of the arm (proprioception), and the motor commands sent to the robot (action). We cannot directly measure the positions of the objects, how they react to being pushed, their speed, etc. Further, in this data, there is no notion of progress or success. Unlike a game of Breakout or hammering a nail, we don’t get a score or an objective. All we have to learn from, when interacting in the real world, is what is provided by our senses, or in this case, the robot’s sensors.</p>
<p>So, what can we learn, when only given our senses? We can learn to <em>predict</em> — what will the world look like, or feel like, if the robots moves its arm in one way versus in another way?</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/predictions_rigid.gif" height="100" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/cloth_crop.gif" height="90" style="margin: 5px;" />     <br /> <i> The robot learns to predict what the future will look like if it moves<br /> its arm in different ways, learning about physics, objects, and itself. </i> </p>
<p>Prediction allows us to learn general things about the world, things like objects and physics. And such general-purpose knowledge is exactly what the Breakout-playing agent is missing. Prediction also allows us to learn from all of the data that we have: a stream of actions and images has a lot of implicit supervision. This is critical because we don’t have a score or reward function. Model-free reinforcement learning systems typically only learn from the supervision provided from the reward function, whereas model-based RL agents utilize the rich information available in the pixels they observe. Now, how do we actually use these predictions? We will discuss this next.</p>
<h2 id="planning-to-perform-human-specified-tasks">Planning to Perform Human-Specified Tasks</h2>
<p>If we have a predictive model of the world, then we can use it to plan to achieve goals. That is, if we understand the consequences of our actions, then we can use that understanding to choose actions that lead to the desired outcome. We use a sampling-based procedure to plan. In particular, we sample many different candidate action sequences, then select the top plans—the actions that are most likely to lead to the desired outcome—and refine our plan iteratively, by resampling from a distribution of actions fitted to the top candidate action sequences. Once we come up with a plan that we like, we then execute the first step of our plan in the real world, observe the next image, and then replan in case something unexpected happened.</p>
<p>A natural question now is—how can a user specify a goal or desired outcome to the robot? We have experimented with a number of different ways to do so. One of the easiest mechanisms that we have found is to simply click on a pixel in the initial image and specify where the object corresponding to that pixel should be moved, by clicking another pixel position. We can also give more than one pair of pixels to specify other desired object motions. While there are types of goals that cannot be expressed in this way (and we have explored more versatile goal specifications, such as goal classifiers), we have found that specifying pixel positions can be used to describe a wide variety of tasks and is remarkably easy to provide. To be clear, these user-provided goal specifications are not used during the data collection, when the robot is interacting with the world—they are only used at test-time per se, when we want the robot to use its predictive model to accomplish a certain goal.</p>
<h2 id="experiments">Experiments</h2>
<p>We experiment with this overall approach on a Sawyer robot, collecting 2 weeks of unsupervised experience. Critically, the only human involvement during training is providing a diverse range of objects for the robot to interact with (swapping out objects periodically) and coding the random robot motions that are used to collect data. This allows us to collect data on multiple robots nearly 24 hours a day, with very little effort. We train a <em>single</em> action-conditioned video prediction model on all of this data, including two camera viewpoints, and use the iterative planning procedure described previously to plan and execute on user-specified tasks.</p>
<p>Since we set out to achieve <em>generality</em>, we evaluate the same predictive model on a wide range of tasks involving objects that the robot has never seen before and goals the robot has not encountered previously.</p>
<p>For example, we ask the robot to fold shorts:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/shorts1.png" height="180" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/shorts2.gif" height="180" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/shorts3.gif" height="180" style="margin: 5px;" />     <br /> <i> Left: The goal is to fold the left side of the shorts. Middle: the robot’s<br /> prediction corresponding to its plan. Right: the robot performs its plan. </i> </p>
<p>Or put an apple on a plate:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/apple1.png" height="180" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/apple2.gif" height="180" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/apple3.gif" height="180" style="margin: 5px;" />     <br /> <i> Left: The goal is to put the apple on the plate. Middle: the robot’s<br /> prediction corresponding to its plan. Right: the robot performs its plan. </i> </p>
<p>Finally, we can also ask the robot to cover a spoon with a towel:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/brown1.png" height="180" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/brown2.gif" height="180" style="margin: 5px;" />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/visual_rl/brown3.gif" height="180" style="margin: 5px;" />     <br /> <i> Left: The goal is to cover the spoon with the towel. Middle: the robot’s<br /> prediction corresponding to its plan. Right: the robot performs its plan. </i> </p>
<p>Interestingly, we find that, even though the model’s predictions are far from perfect, it can still use them to effectively accomplish the specified goal.</p>
<h2 id="related-work">Related Work</h2>
<p><a href="http://mlg.eng.cam.ac.uk/pub/pdf/DeiRas11.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">There</a> <a href="http://www.ece.utah.edu/~bodson/acscr/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">have</a> <a href="https://arxiv.org/abs/1511.09249" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">been</a> <a href="https://cs.stanford.edu/people/asaxena/papers/deepmpc_rss2015.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">many</a> <a href="https://arxiv.org/abs/1708.02596" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">prior</a> <a href="https://arxiv.org/abs/1805.12114" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">works</a> that approach the problem of model-based reinforcement learning (RL), i.e. learning a predictive model, and then using this model to act or using it to learn a policy. Many of such prior works have focused on settings where the the positions of objects or other task-relevant information can be accessed directly—rather than through images or other raw sensor observations. Having this low-dimensional state representation is a strong assumption that is often impossible to fulfill in the real world . Model-based RL methods that directly operate on raw image frames have not been studied as extensively.  Several algorithms have been proposed for simple, <a href="https://arxiv.org/abs/1502.02251" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">synthetic</a> <a href="https://arxiv.org/abs/1506.07365" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">images</a> and <a href="https://arxiv.org/abs/1611.01779" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">video</a> <a href="https://worldmodels.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">game</a> <a href="https://arxiv.org/abs/1507.08750" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">environments</a>, which have focused on a fixed set of objects and tasks. <a href="https://arxiv.org/abs/1509.06113" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Other</a> <a href="https://arxiv.org/abs/1710.00489" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">work</a> <a href="https://arxiv.org/abs/1808.09105" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">has</a> studied model-based RL in the real world, again focusing on individual skills.</p>
<p>A number of recent works have studied self-supervised robotic learning, where large-scale unattended data collection is used to learn individual skills such as grasping (e.g. <a href="https://arxiv.org/abs/1509.06825" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">see</a> <a href="https://arxiv.org/abs/1603.02199" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">these</a> <a href="https://arxiv.org/abs/1710.05512" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">works</a>), <a href="https://arxiv.org/abs/1803.09956" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">push-grasp synergies</a>, or <a href="https://arxiv.org/abs/1702.01182" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">obstacle</a> <a href="https://arxiv.org/abs/1704.05588" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">avoidance</a>. Our approach is also fully self-supervised; in contrast with these approaches, we learn a predictive model that is goal-agnostic and can be used to perform a variety of manipulation skills.</p>
<h2 id="discussion">Discussion</h2>
<p>Generalization to many distinct tasks in visually diverse settings is arguably one of the biggest challenges in reinforcement learning and robotics research today. Deep learning has greatly reduced the amount of task-specific engineering needed to deploy an algorithm; however, prior methods typically require extensive amounts of supervised experience or focus on mastery of individual tasks. Our results suggest that our approach can generalize to a wide range of tasks and objects, including those never seen previously. The generality of the model is the result of large-scale self-supervised learning from interaction. We believe the results represent a significant step forward in terms of <em>generality</em> of tasks achieved by a single robotic reinforcement learning system.</p>
<hr />
<p><strong>The work in this post is based on the following paper</strong>:</p>
<ul>
<li><strong>Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control</strong><br /> <a href="https://febert.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Frederik Ebert</a>*, <a href="http://people.eecs.berkeley.edu/~cbfinn/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Chelsea Finn</a>*, Sudeep Dasari, Annie Xie, <a href="https://people.eecs.berkeley.edu/~alexlee_gk/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Alex Lee</a>, <a href="http://people.eecs.berkeley.edu/~svlevine" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sergey Levine</a><br /> <a href="https://sites.google.com/view/visualforesight" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project webpage</a></li>
</ul>
<p><strong>The above paper is an extended version of the following four papers, and builds upon the fifth paper</strong>:</p>
<ul>
<li>
<p><strong><a href="https://arxiv.org/abs/1810.03043" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Robustness via Retrying: Closed-Loop Robotic Manipulation via Self-Supervised Learning</a></strong><br /> <a href="https://febert.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Frederik Ebert</a>, Sudeep Dasari, <a href="https://people.eecs.berkeley.edu/~alexlee_gk/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Alex Lee</a>, <a href="http://people.eecs.berkeley.edu/~svlevine" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sergey Levine</a>, <a href="http://people.eecs.berkeley.edu/~cbfinn/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Chelsea Finn</a><br /> Conference on Robot Learning (CoRL), 2018</p>
</li>
<li>
<p><strong><a href="https://arxiv.org/abs/1810.00482" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Few-Shot Goal Inference for Visuomotor Learning and Planning</a></strong><br /> Annie Xie, <a href="http://people.eecs.berkeley.edu/~avisingh/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Avi Singh</a>, <a href="http://people.eecs.berkeley.edu/~svlevine" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sergey Levine</a>, <a href="http://people.eecs.berkeley.edu/~cbfinn/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Chelsea Finn</a><br /> Conference on Robot Learning (CoRL), 2018</p>
</li>
<li>
<p><strong><a href="https://arxiv.org/abs/1710.05268" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Self-Supervised Visual Planning with Temporal Skip Connections</a></strong><br /> <a href="https://febert.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Frederik Ebert</a>, <a href="http://people.eecs.berkeley.edu/~cbfinn/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Chelsea Finn</a>, <a href="https://people.eecs.berkeley.edu/~alexlee_gk/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Alex Lee</a>, <a href="http://people.eecs.berkeley.edu/~svlevine" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sergey Levine</a><br /> Conference on Robot Learning (CoRL), 2017</p>
</li>
<li>
<p><strong><a href="https://arxiv.org/abs/1610.00696" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Deep Visual Foresight for Planning Robot Motion</a></strong><br /> <a href="http://people.eecs.berkeley.edu/~cbfinn/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Chelsea Finn</a> &amp; <a href="http://people.eecs.berkeley.edu/~svlevine" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sergey Levine</a><br /> International Conference on Robotics and Automation (ICRA), 2017</p>
</li>
<li>
<p><strong><a href="https://arxiv.org/abs/1605.07157" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Unsupervised Learning for Physical Interaction via Video Prediction</a></strong><br /> <a href="http://people.eecs.berkeley.edu/~cbfinn/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Chelsea Finn</a>, Ian Goodfellow, <a href="http://people.eecs.berkeley.edu/~svlevine" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sergey Levine</a><br /> Neural Information Processing Systems (NeurIPS), 2016</p>
</li>
</ul>
<p>This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>AdaSearch: A successive elimination approach to adaptive search</title>
		<link>https://robohub.org/adasearch-a-successive-elimination-approach-to-adaptive-search/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 14 Nov 2018 23:12:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/adasearch-a-successive-elimination-approach-to-adaptive-search/</guid>

					<description><![CDATA[  




In many tasks in machine learning, it is common to want to answer questions
given fixed, pre-collected datasets. In some applications, however, we are not
given data a priori; instead, we must collect the data we require to answer the
questions...]]></description>
										<content:encoded><![CDATA[<p><strong>By Esther Rolf∗, David Fridovich-Keil∗, and Max Simchowitz</strong></p>
<div class="keep-aspect"><iframe title="AdaSearch (5 min)" width="500" height="281" src="https://www.youtube-nocookie.com/embed/a4SPB3VugFI?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></div>
<p></p>
<p>In many tasks in machine learning, it is common to want to answer questions given fixed, pre-collected datasets. In some applications, however, we are not given data a <em>priori</em>; instead, we must collect the data we require to answer the questions of interest.<span id="more-112613"></span></p>
<p>This situation arises, for example, in environmental contaminant monitoring and census-style surveys. Collecting the data ourselves allows us to focus our attention on just the most relevant sources of information. However, determining which of these sources of information will yield useful measurements can be difficult. Furthermore, when data is collected by a physical agent (e.g. robot, satellite, human, etc.) we must plan our measurements so as to reduce costs associated with the motion of the agent over time. We call this abstract problem <em>embodied adaptive sensing</em>.</p>
<p>We introduce a new approach to the embodied adaptive sensing problem, in which a robot  must traverse its environment to identify locations or items of interest. Adaptive sensing encompasses many well-studied problems in robotics, including the rapid identification of accidental contamination leaks and radioactive sources, and finding individuals in search and rescue missions. In such settings, it is often critical to devise a sensing trajectory that returns a correct solution as quickly as possible.</p>
<p>  <!--more-->  </p>
<p>We focus on the problem of radioactive source-seeking (RSS), in which a UAV must identify the $k$-largest radioactive emitters in its environment, where $k$ is a user-defined parameter. RSS is a particularly interesting instance of the adaptive sensing problem, due both to the challenges posed by the highly heterogeneous background noise, as well as to the existence of a well-characterized sensor model amenable to the construction of statistical confidence intervals.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adasearch/adasearch.png" width="600" alt="..." />      </p>
<p>We introduce AdaSearch, a successive-elimination framework for general adaptive sensing problems, and demonstrate it within the context of radioactive source seeking. AdaSearch explicitly maintains confidence intervals over the emissions rate at each point in the environment. Using these confidence intervals, the algorithm iteratively identifies a set of candidate points likely to be among the top emitters, and eliminates other points.</p>
<h1 id="embodied-search-as-a-multiple-hypothesis-testing-scenario">Embodied Search as a Multiple Hypothesis Testing Scenario</h1>
<p>Traditionally, the robotics community has conceived of embodied search as a continuous motion planning problem, where the robot must balance exploring its environment with selecting efficient trajectories. This has motivated approaches where both trajectory optimization and exploration are combined into a single objective, which can be optimized using receding horizon control (<a href="https://ieeexplore.ieee.org/abstract/document/5350445" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Hoffman and Tomlin</a>, <a href="http://personal.stevens.edu/~benglot/Bai_Wang_Chen_Englot_IROS2016_AcceptedVersion.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Bai et al.</a>, <a href="https://ieeexplore.ieee.org/document/6385653" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Marchant and Ramos</a>). Instead, we consider an alternate approach in which we formulate the problem as one of sequential best-action identification via hypothesis testing.</p>
<p>In sequential hypothesis testing, the goal is to reach conclusions on many separate questions, by iteratively collecting data. An agent is given a set of $N$ measurement actions, each of which yields observations according to a distinct, fixed distribution.</p>
<p>The agent’s goal is to learn some prespecified property of these $N$ observation distributions. For example, in a statistical “A/B test,” a measurement action corresponds to showing a new customer either product A or product B, and recording their assessment of that product.  Here, $N=2$ because there are just two actions, showing product A and showing product B. The property of interest is which product is preferred on average (B in the illustration below).  As we collect measurements on preferences, we keep track of sample averages, as well as confidence intervals around them, described by a lower confidence bound (LCB) and an upper confidence bound (UCB) for each product.  As we collect more measurements, we become more confident in our estimate of preference for each individual product, and therefore our ranking between products. This suggests a condition for concluding that product B is preferred to product A: <em>If the LCB for product B is greater than the UCB for product A, then we can conclude that with high probability, B is prefered to A, on average.</em></p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adasearch/abtesting_cropped.gif" width="600" alt="..." />      </p>
<p>In the context of environmental sensing, each action may correspond to taking a sensor reading from a given position and orientation. Typically, the agent wishes to know which single measurement action yields observations with the greatest mean observed signal, or which set of $k$ actions together have the greatest mean observations. To do so, the agent may choose actions <em>sequentially</em>, using previously measured observations to favor future actions which are most informative for discerning the actions with largest mean observations.</p>
<p>At first glance, sequential best-action identification may seem like too abstract a framework to be useful in mobile, embodied sensing agents. Indeed, the agent can choose any arbitrary sequence of measurement actions, without considering the potential costs—such as movement time—associated with changing actions. However, the abstract nature of sequential best-action identification is also its most formidable strength. By formulating the embodied search problem in precise statistical language, we develop actionable confidence intervals about the observation means associated with each sensing action, and determine the set of all actions which still need to be taken before the points of interest can confidently determined.</p>
<p>Our proposed approach to embodied search, AdaSearch, uses confidence intervals from sequential best-action identification and a global trajectory planning heuristic to both achieve asymptotically optimal measurement complexity, and effectively amortize movement costs.</p>
<h1 id="radioactive-source-seeking">Radioactive Source Seeking</h1>
<p>For concreteness, we will present AdaSearch in the context of the radioactive source seeking problem with a single source. We model the environment as a planar grid, as in the depiction below. There is exactly one high-intensity radioactive point source (red dot). However, locating this source is difficult because sensor measurements are corrupted by the background radiation (pink dots). Sensor measurements are obtained by flying a quadrotor equipped with a radiation sensor above the grid. The goal is to devise a sequence of trajectories so that the measurements obtained from the onboard sensor allow us to disambiguate the radioactive point source from the background radiation sources, as quickly as possible.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adasearch/adasearch_clip_intro.gif" width="600" alt="..." />      </p>
<h2 id="adasearch">AdaSearch</h2>
<p>Our algorithm, AdaSearch, combines a global-coverage planning approach with an adaptive sensing rule based on hypothesis testing to define these trajectories. In the first pass through the grid, we sample uniformly over the environment.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adasearch/adasearch_pass1_clip.gif" width="600" alt="..." />      </p>
<p>After observing the measurements during the first pass, we can eliminate some regions from consideration. Points are eliminated if the upper bound of our estimated confidence interval around their mean is smaller than the largest lower bound of any interval. This means that with high probability, they are not the source that we’re looking for.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adasearch/adasearch_pass2_clip.gif" width="600" alt="..." />      </p>
<p>In the next round, AdaSearch focuses on sampling the remaining points (teal squares) more carefully, because they are still potential source locations.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adasearch/adasearch_pass3_clip.gif" width="600" alt="..." />      </p>
<p>This process continues, and with each round the set of candidate source locations shrinks until only a single point remains. AdaSearch returns this point (enlarged red point) as the radioactive point source we were searching for.</p>
<p>Due to the crisp statistical formulation of confidence, we can be sure that under known sensing models, AdaSearch returns the correct source with high probability. We ensure a certain level of confidence in this probabilistic guarantee by fixing the width (in standard deviations) of the confidence bounds around each individual region, throughout the course of the algorithm. Furthermore, AdaSearch comes with environment-specific runtime guarantees, as we describe in detail <a href="https://arxiv.org/abs/1809.10611" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">in our paper</a>.</p>
<h2 id="baselines">Baselines</h2>
<p>Perhaps the most popular approach for general adaptive search problems is information maximization (<a href="https://ieeexplore.ieee.org/abstract/document/1041446" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Bourgault et al.</a>). Information maximization methods collect measurements in locations deemed promising according to an information theoretic criterion, and follow a receding horizon strategy to plan trajectories. We compare AdaSearch to a version of information maximization tailored to radiation detection: InfoMax.</p>
<p>Unfortunately, for large search spaces, the real-time computational constraints of this approach necessitate approximations such as limits on planning horizon and trajectory parameterization. These approximations may cause the algorithm to be excessively greedy and spend too much time tracking down false leads.</p>
<p>To disambiguate between the effects of our statistical confidence intervals and global planning heuristic (vs. InfoMax’s information metric and receding horizon planning), we implement as a simple global planning approach, NaiveSearch, as a second baseline. This approach samples the grid uniformly, spending an equal amount of time at each grid cell.</p>
<h2 id="results">Results</h2>
<p>We implemented all three algorithms and simulated their performance on ten randomized instantiations of the problem on a 64 by 64 meter grid, at 4 meter resolution, using realistic quadrotor dynamics and simulated radiation sensor readings.</p>
<p>In our experiments, we observe that AdaSearch usually finishes faster than NaiveSearch and InfoMax. As we increase the maximum background radiation level, the ratio of AdaSearch’s run time to NaiveSearch’s runtime continues to improve, which matches the theoretical bounds given in <a href="https://arxiv.org/abs/1809.10611" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">the full paper</a>.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/adasearch/runtime.png" alt="..." />      </p>
<p>The increase in performance between AdaSearch  and NaiveSearch suggests that adaptivity does confer advantages over non-adaptive methods. Even more strikingly, and somewhat unexpectedly, even NaiveSearch tends to outperform InfoMax in this problem setting. This suggests that the locally greedy nature of receding horizon control in InfoMax is indeed hurting its performance. AdaSearch, by contrast, gracefully blends adaptive strategies with global coverage guarantees.</p>
<h1 id="adasearch-more-generally">AdaSearch more generally</h1>
<p>The successful demonstration of AdaSearch operating onboard a UAV in the context of finding radioactive sources prompts us to ask, <em>in what more general problem settings will AdaSearch also perform well</em>? As it turns out, the core algorithm applies more broadly, even to non-robotic embodied sensing problems.</p>
<p>For example, consider the problem of planning a pilot program at 10 out of 100 medical clinics spread across a region. We might wish to establish these programs in the locations with the highest rates of a particular rare disease by conducting surveys across potential clinic locations to assess rates of illness in each region. This is an embodied sensing problem, as diagnoses are made in person. Resources are limited in terms of the number of human surveyors, and there are physical constraints both on the time required to survey a group of people, and on the travel time between towns.</p>
<p>A survey planner could use AdaSearch to guide the decisions of how long to spend in each potential clinic location counting new cases of the disease before moving on to the next, and to trade-off the travel time of returning to collect more data from a certain town with spending extra time at the town in the first place.</p>
<p>In general, AdaSearch is expected to perform well when we think that measurements are noisy enough to warrant multiple passes through space when collecting data. Radioactive gamma ray emissions, as well as occurances of a rare disease, can be modeled as Poisson distributed random variables, where the variance scales with the mean.  AdaSearch easily adapts to different noise models (e.g. Gaussian), which might arise with different applications. So long as we can calculate or bound the appropriate confidence intervals, AdaSearch guarantees an efficient traversal of the region to find the points of interest.</p>
<p>For more information about AdaSearch, please see full video at the top of the page, and the full text of the paper at: <a href="https://arxiv.org/abs/1809.10611" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">https://arxiv.org/abs/1809.10611</a>. This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Drilling down on depth sensing and deep learning</title>
		<link>https://robohub.org/drilling-down-on-depth-sensing-and-deep-learning/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 24 Oct 2018 20:54:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/drilling-down-on-depth-sensing-and-deep-learning/</guid>

					<description><![CDATA[    
    
    
    
    
    
    
    
    

Top left: image of a 3D cube. Top right: example depth image, with darker points
representing areas closer to the camera (source: Wikipedia). Next two
rows: examples of depth and RGB image pairs for graspi...]]></description>
										<content:encoded><![CDATA[<p><strong>By Daniel Seita, Jeff Mahler, Mike Danielczuk, Matthew Matl, and Ken Goldberg</strong></p>
<p>This post explores two independent innovations and the potential for combining them in robotics. Two years before the <a href="https://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">AlexNet results</a> on <a href="http://image-net.org/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ImageNet</a> were released in 2012, Microsoft rolled out the Kinect for the X-Box. This class of low-cost depth sensors emerged just as Deep Learning boosted Artificial Intelligence by accelerating performance of hyper-parametric function approximators leading to surprising advances in <a href="https://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">image classification</a>, <a href="https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/38131.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">speech recognition</a>, and <a href="https://papers.nips.cc/paper/5346-sequence-to-sequence-learning-with-neural-networks.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">language translation</a>.<span id="more-111091"></span></p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/depth_sensing/two_cubes_side_by_side_v01.png" width="500" alt="..." />     <br />     <!--     <img decoding="async" src="http://bair.berkeley.edu/static/blog/depth_sensing/ms_mike.png"          width="700"          alt="..."/>     <br />     -->     <img decoding="async" src="http://bair.berkeley.edu/static/blog/depth_sensing/Dex-Net-Depth-v2.jpg" width="700" alt="..." />     <br />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/depth_sensing/dex-net-rgb-v2.png" width="700" alt="..." />     <br />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/depth_sensing/ms_daniel.png" width="700" alt="..." />     <br /> <i> Top left: image of a 3D cube. Top right: example depth image, with darker points representing areas closer to the camera (<a href="https://en.wikipedia.org/wiki/Depth_map" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">source: Wikipedia</a>). Next two rows: examples of depth and RGB image pairs for grasping objects in a bin.  Last two rows: similar examples for bed-making. </i> </p>
<p> Today, Deep Learning is also showing promise for end-to-end learning of <a href="https://www.nature.com/articles/nature14236" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">playing video games</a> and <a href="https://arxiv.org/abs/1504.00702" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">performing robotic manipulation</a> tasks.</p>
<p>For robot perception, <a href="http://cs231n.github.io/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">convolutional neural networks</a> (CNNs), such as <a href="https://arxiv.org/abs/1409.1556" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">VGG</a> or <a href="https://arxiv.org/abs/1512.03385" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ResNet</a>, with three RGB color channels have become standard. For robotics and computer vision tasks, it is common to borrow one of these architectures (along with pre-trained weights) and then to <a href="http://cs231n.github.io/transfer-learning/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">perform transfer learning or fine-tuning</a> on task-specific data. But in some tasks, knowing the colors in an image may provide only limited benefits. Consider training a robot to grasp novel, previously unseen objects.  It may be more important to understand the geometry of the environment rather than colors and textures. The physical process of manipulation — controlling one or more objects by applying forces through contact — depends on object geometry, pose, and other factors which are largely color-invariant. When you manipulate a pen with your hand, for instance, you can often move it seamlessly without looking at the actual pen, so long as you have a good understanding of the location and orientation of contact points.  Thus, before proceeding, one might ask: <em>does it makes sense to use color images?</em></p>
<p>There is an alternative: <em>depth images</em>. These are single-channel grayscale images that measure depth values from the camera, and give us invariance to the colors of objects within an image. We can also use depth to “filter” points beyond a certain distance which can remove background noise, as we demonstrate later with robot bed-making. Examples of paired depth and real images are shown above.</p>
<p>In this post, we consider the potential for combining depth images and deep learning in the context of three ongoing projects in the <a href="http://autolab.berkeley.edu/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">UC Berkeley AUTOLab</a>: Dex-Net for robot grasping, segmenting objects in heaps, and robot bed-making.</p>
<p>  <!--more-->  </p>
<h1 id="sensing-depth">Sensing Depth</h1>
<p>Depth images encode distance (e.g., in millimeters) of surfaces in a scene relative to a particular viewpoint. We provide an example in the image at the top of this post. On the top left is an RGB image of a 3D cube structure, which has points located at a variety of distances from the camera. To the top right is one representation of a depth image, with darker points representing closer surfaces, though it is also valid to use other representations, such as using darker points for <em>farther</em> areas, or to use depth with respect to a different origin. For additional background on how depth images can be created, <a href="https://blog.cometlabs.io/depth-sensors-are-the-key-to-unlocking-next-level-computer-vision-applications-3499533d3246" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">check out this blog post by the Comet Labs Research Team</a>.</p>
<h2 id="recent-advancements-in-depth-sensing">Recent Advancements in Depth Sensing</h2>
<p>Recently, there have been a number of advancements in depth sensing which have occurred in parallel with improvements in computer vision and deep learning.</p>
<p>Classically, depth sensing involved matching pairs of points between aligned RGB images from two different cameras, and then using the resulting disparity map to obtain the depth of objects in the environment.</p>
<p>The depth sensors we commonly use today are <em>structured light sensors</em>, which project a known pattern into the scene using a non-visible wavelength. The Kinect innovation in particular was to project a known pattern from an infrared (IR) projector and image that pattern with a single IR camera. Since light travels in straight lines, a virtual IR camera placed at the projector would always capture the same image of the pattern. Therefore, the image pattern from the real IR camera can be matched against a pre-saved “template” image to find correspondences. This can be done quickly on embedded hardware.</p>
<p>Another approach to depth sensing is <a href="https://en.wikipedia.org/wiki/Lidar" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">LIDAR</a>, an older technique which is commonly used for surveying land and terrain, and has recently been applied for some self-driving cars. LIDAR, while generally providing higher-quality depth maps than Kinect, is slower and more expensive due to the need to scan lasers.</p>
<p>In sum, the Kinect is a consumer-grade RGB-D system that captures RGB images along with per-pixel depth values directly with the hardware, and is faster and cheaper (without sacrificing too much accuracy) than prior solutions. Nowadays, many robots available today for research and industrial purposes, such as the <a href="https://fetchrobotics.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Fetch Robot</a> and the <a href="https://www.toyota-global.com/innovation/partner_robot/robot/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Toyota Human Support Robot</a>, come equipped with similar built-in depth sensing cameras. Future advancements in depth sensing for robots may come from improvements in existing cameras such as <a href="https://www.intel.com/content/www/us/en/architecture-and-technology/realsense-overview.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Intel’s RealSense</a>, or from newer technologies introduced by companies such as <a href="https://www.photoneo.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Photoneo</a>.</p>
<h2 id="prior-research-using-depth-images">Prior Research Using Depth Images</h2>
<p>The availability of depth sensing in robotics hardware has allowed depth images to be used for <a href="http://www2.informatik.uni-freiburg.de/~hornunga/pub/maier12humanoids.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">real-time navigation</a> (Maier et al., 2012), for <a href="https://ieeexplore.ieee.org/document/6162880/authors#authors" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">real-time mapping and tracking</a> (Newcombe et al., 2011), and for <a href="https://homes.cs.washington.edu/~xren/publication/henry-ijrr12-rgbd-mapping.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">modeling indoor environments</a> (Henry et al., 2012). Since depth allows robots to understand how far they are from obstacles, it enables them to locate and avoid them during navigation.</p>
<p>Depth images have additionally been used to <a href="http://www.robotics.stanford.edu/~koller/Papers/Plagemann+al:ICRA10.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">detect, identify and localize body parts of humans in real time</a> (Plagemann*, Ganapathi*, et al., 2010) with high reliability on real gaming systems (e.g., the Xbox One). Depth could remove or mitigate sources of ambiguity, such as lighting and the wide variety of human appearances and clothing. Other recent work uses simulated depth images to <a href="https://arxiv.org/abs/1706.04652" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">develop closed-loop policies to guide a robot arm towards an object</a> (Viereck et al., 2017). In their case, the advantage of depth images was that large datasets could be rapidly generated in simulation, and the depth images were simulated relatively accurately using ray tracing.</p>
<p>These results suggest that for some tasks, depth images can encode a sufficient amount of useful information and color invariance can be beneficial. We describe three such cases below.</p>
<h1 id="example-1--robot-grasping">Example 1:  Robot Grasping</h1>
<p>Universal picking – grasping a large variety of previously unseen objects – remains a Grand Challenge for robotics.  Although many researchers (e.g., <a href="https://arxiv.org/abs/1509.06825" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Pinto and Gupta, 2016</a>) use RGB images, their systems need many months of training time with robots physically executing grasps. A key advantage of using 3D object meshes is that one can synthesize accurate depth images via rendering techniques, which use geometry and camera projection (<a href="https://arxiv.org/abs/1608.02239" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Johns et al., 2016</a>, <a href="https://arxiv.org/abs/1706.04652" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Viereck et al., 2017</a>).</p>
<p>Our <a href="https://berkeleyautomation.github.io/dex-net/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Dexterity Network (Dex-Net)</a> is an ongoing research project in the AUTOLab that encompasses algorithms, code, and datasets for training robot grasping policies using a combination of large synthetic datasets, analytic robustness models, stochastic sampling, and deep learning techniques. Dex-Net introduced domain randomization in the context of grasping, focusing on grasping complex objects with a simple gripper in contrast to <a href="https://blog.openai.com/learning-dexterity/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recent work from OpenAI</a> showing the value of domain randomization for grasping simple objects with a complex gripper.  In <a href="https://bair.berkeley.edu/blog/2017/06/27/dexnet-2.0/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">a prior BAIR Blog post</a>, we presented a dataset with 6.7 million samples in it, which was used to train a grasp quality model. Here, we expand the discussion with a focus on depth images.</p>
<h2 id="dataset-and-depth-images">Dataset and Depth Images</h2>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/depth_sensing/data_dexnet.png" alt="..." />     <br /> <i> The dataset generation process for Dex-Net. First, a large number of object mesh models are generated and augmented from a variety of sources. For each model, multiple parallel-jaw grasps are sampled for it. For each object and grasp combination, we compute the robustness and generate a simulated depth image. Robustness is computed by estimating the probability of grasp success over a stochastic distribution on pose, friction, mass, and external forces (e.g., gravity direction) with Monte-Carlo Integration. To the right, we show samples of positive and negative (success vs failure) grasp attempts, and show the images that the network sees; the red grasp overlays are only for visualization purposes. (Open in a new window to enlarge.) </i> </p>
<p>We recently extended Dex-Net to automatically generate a modified synthetic dataset of grasps on object meshes. Grasps are specified as the planar position, angle, and depth of a gripper relative to an RGB-D sensor. We present an overview of the data formation process in the figure above. Our overall goal is to train a deep network that can detect whether a grasp attempt on some (singulated) object, represented in a depth image, will succeed.</p>
<h2 id="training-a-gq-cnn">Training a GQ-CNN</h2>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/depth_sensing/network_dexnet.png" alt="..." />     <br /> <i> The Grasp Quality CNN architecture. A grasp candidate image (shown to the left) is processed and aligned based on the angle and center of the grasp, and a corresponding 96&#215;96 depth image (labeled &#8220;Aligned Image&#8221;) is passed as input, along with the height $z$, to predict grasp robustness. </i> </p>
<p>The simulated dataset is used to train a Grasp Quality Convolutional Neural Network (GQ-CNN) to determine how likely a grasp attempt will succeed. One can use this GQ-CNN in a policy. For example, a policy could sample various grasps and feed each through the GQ-CNN, pick the one with the highest grasp success probability, and then execute its corresponding open-loop trajectory. For an overview of our results, please see <a href="https://bair.berkeley.edu/blog/2017/06/27/dexnet-2.0/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our prior BAIR Blog post</a>.</p>
<p>In 2017, <a href="http://proceedings.mlr.press/v78/mahler17a/mahler17a.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Dex-Net was extended to bin-picking</a>, which involves iteratively grasping objects from heaps. We modeled bin-picking as a Partially Observed Markov Decision Process, and generated object heaps via simulation. Due to the simulation, we were able to obtain full knowledge of the object poses, and used an algorithmic supervisor to perform demonstrations of the task. We then fine-tuned a GQ-CNN and performed imitation learning on the supervisor’s policy. Using the resulting learned policy on a physical <a href="https://new.abb.com/products/robotics/industrial-robots/yumi" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ABB YuMi robot</a>, we were able to clear heaps of 10 objects in under three minutes using only information from the depth cameras.</p>
<p>Below, we show examples of real and simulated depth images which show grasps from the Dex-Net system in a setup with multiple objects in a bin.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/depth_sensing/blog_post_sim_vs_real-01.png" alt="..." />     <br /> <i> Top row: real depth images taken from the camera mounted over our ABB YuMi robot. Bottom row: simulated depth images from Dex-Net. The red overlays indicate the grasp attempt. </i> </p>
<h1 id="example-2--segmenting-objects-in-bins">Example 2:  Segmenting Objects in Bins</h1>
<p>Instance segmentation is the task of determining which pixels in an image belong to which object, while also separating instances of the same <em>class</em>. Instance segmentation is widely used for robot perception; for example, as the initial step in a robotic perception pipeline for grasping objects cluttered in a bin, where the robot first segments the image to localize the target object to grasp before executing a grasping policy.</p>
<p>Prior research in computer vision has demonstrated that <a href="https://arxiv.org/abs/1703.06870" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Mask R-CNN</a> can be trained to segment objects in RGB images, but this training requires massive hand-labeled datasets of real RGB images. In addition, images used for training Mask R-CNN tend to represent natural scenes with limited numbers of object classes. Thus, pretrained Mask R-CNN networks may not perform well on a task such as segmenting arbitrary objects in a warehouse bin, and fine-tuning would require knowledge and hand-labeled examples of each object. If we relax our requirement that we predict each object’s class in addition to its mask, we can predict masks for a larger set of object classes, and object <em>geometries</em> become more influential than object <em>identities</em>.</p>
<h2 id="dataset-and-depth-images-1">Dataset and Depth Images</h2>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/depth_sensing/segmentation_data.png" alt="..." />     <br /> <i> Our dataset formation process. Left: we sample 3D object models similar to those used in Dex-Net. These are shuffled and dropped into an object heap, either through simulation or through physical experiments. The corresponding depth images are created, along with object masks for training and ground-truth evaluation. </i> </p>
<p>For geometry-based segmentation, we can use simulation and rendering techniques to automate the process of collecting large and diverse training datasets of labeled depth images, as shown in the figure above. We hypothesize that these depth images may contain enough information about <em>segmentation cues</em>, since discontinuities are indicative of the “pixel borders” of objects. Our simulated dataset of 50K depth images was generated by sampling several 3D objects out of 1600 models, and dropping them into a bin via PyBullet simulation. Since the object models are known, we can automatically generate accurate depth images along with ground-truth masks. Using this dataset, we trained a version of Mask R-CNN, which we call <strong>SD Mask R-CNN</strong>, <em>only</em> on synthetic depth data.</p>
<h2 id="segmentation-results-on-real-images">Segmentation Results on Real Images</h2>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/depth_sensing/results_segmentation.png" alt="..." />     <br /> <i> The results suggest that our SD Mask R-CNN can accurately segment despite not being trained on any real images. We show an example bin picking setup, the depth image, the ground truth segmentation, two baselines, and our method. The two rows represent the same bin-picking setup, but with two sensors: high resolution (top) and low resolution (bottom). (Open in a new window to enlarge.) </i> </p>
<p>Our proposed SD Mask R-CNN outperforms point cloud segmentation and fine-tuned Mask R-CNN on a dataset of real images <em>despite not being trained on any real images</em>. An example of segmentation results and other related images are shown above. Importantly, the objects used in creating the hand-labeled dataset of real images were not chosen from the training distribution of SD Mask R-CNN; in fact, they were common household items for which we do not have 3D models. Thus, SD Mask R-CNN can predict masks for previously unseen objects. Moreover, we find that we can reduce the size of the backbone network of Mask R-CNN (e.g., from ResNet-101 to ResNet-35) for depth images as compared to training with color images.</p>
<p>Segmenting objects as “object” or “background” allows for decoupling of the classification and segmentation stages; we found that for a set of ten objects, we could train a VGG classifier in less than ten minutes that could achieve over 95% classification accuracy. These results suggest that SD Mask R-CNN could be used in tandem with a classification network, which could easily be retrained for each set of objects used.</p>
<p>Overall, our segmentation results suggest three main benefits of using depth over RGB images:</p>
<ol>
<li>depth information may encode the geometric cues necessary to separate object instances both from each other and the background of the image,</li>
<li>synthetic depth images can be easily and rapidly generated, and training on them can effectively transfer to real images,</li>
<li>a network trained using depth images can potentially generalize better to previously unseen objects, as geometric cues can be more consistent across objects.</li>
</ol>
<h1 id="example-3--robot-bed-making">Example 3:  Robot Bed-Making</h1>
<p>Bed-Making is a task we believe could be well-suited for home robotics since it is tolerant to error, not time critical, and rarely enjoyed by humans. <a href="https://bair.berkeley.edu/blog/2017/10/26/dart/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">We introduced the bed-making task in an earlier blog post</a> and explored it with RGB images as a sequential decision problem with noise injection applied for better imitation learning. <a href="https://arxiv.org/abs/1809.09810" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">In our recent preprint</a>, we used depth sensing to extend this project to explore transfer between blankets of different colors and textures and between robots.</p>
<h2 id="dataset-and-depth-images-2">Dataset and Depth Images</h2>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/depth_sensing/init_states_v02.png" alt="..." />     <br /> <i> Examples of initial states for the bed-making task. The first four columns show samples in the training data. The last two demonstrate the Cal and Teal blankets that we used to test generalization to different blankets. (Open in a new window to enlarge.) </i> </p>
<p>We framed the bed-making task as one of detecting corners of a blanket, so that a mobile home robot such as the Fetch or the HSR, can grasp and pull the blanket to a corner of the bed frame to maximize blanket coverage. Our starting hypothesis was that depth images contained enough information about the geometry of blanket corners to allow for reliable bed-making.</p>
<p>To collect training data, we use white blankets with marked red corners, as shown in the above image, so that we can automatically detect a corner and thus a grasping target. We repeatedly toss blankets on the bed surface and collect RGB and depth images from the robot’s onboard RGB-D sensors.</p>
<p>We next train a deep convolutional neural network to detect corners from <em>depth images only</em>, with the hope that the network will generalize to detecting corners from depth images of different blankets. Our deep network utilizes pre-trained weights from <a href="https://arxiv.org/abs/1506.02640" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">YOLO, a fast real-time object detector</a>, because the task of finding a grasping point is similar to detection. We then add several layers after this, which we train with our dataset of 2018 depth images (and yes, 2018 is just a coincidence). Our results indicate that using pre-trained weights is beneficial <em>despite the depth versus RGB mismatch</em>; the pre-trained weights from YOLO were obtained by training on RGB images.</p>
<p>Another advantage of depth images is that it lets us remove sources of distraction. For example, we want the robots to grasp blanket corners. These are not located in areas far beyond the top surface of the bed. Thus, we can “black out” regions beyond a validation-tuned depth value (we used 1.4 meters as the cutoff) before scaling pixels within $[0,255]$. <a href="https://gist.github.com/DanielTakeshi/8fe06f1ea1985bb9246abd5fd21c330e" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">We provide a simple script here</a> that one can use for processing depth images.</p>
<h2 id="corner-detection-results">Corner Detection Results</h2>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/depth_sensing/example_teal_generalization_v02.png" width="600" alt="..." />     <br /> <i> Visualization of a bed-making rollout with an emphasis on corner detection. In the first row, the robot&#8217;s trained grasp policy correctly identifies the corner (top left) and the resulting situation after the grasp and pull is in the top right. On the other side, the policy again detects the corner well (bottom left) and the resulting grasp and pull is shown in the bottom right. </i> </p>
<p>We deployed our trained grasping policy and found that in terms of blanket coverage, it significantly outperformed a non-learning baseline policy, and was nearly as good as a human supervisor. While our metric here is blanket coverage rather than detecting corners, accurate detection is strongly correlated with higher coverage.</p>
<p>In the above image, we show the corner predictions on a teal blanket with the red cross hair. The grasping network was not trained on teal blankets, and only saw the depth images, but nonetheless is able to detect corners accurately since the test-time depth images look similar to depth images from training. After the robot moves to the other side of the bed to attempt another grasp, it again does an excellent job in detecting the nearest corner. We tried using RGB-trained grasping policies, but these did not perform well since the original RGB trained policy was only on white blankets, and we would need far more blankets and training data to generalize across blanket colors.</p>
<h1 id="depth-matters">Depth Matters</h1>
<p>Our results in these projects suggest that depth maps contain sufficient clues for the tasks of determining grasp points, segmenting images, and detecting corners of deformable objects. We conjecture that, as the quality of depth cameras improves in tandem with reduction in costs, depth images will be an increasingly important modality for robotics.   It is far easier to synthesize training examples with depth images, color-invariance results naturally, and background noise can be easily filtered (as we demonstrate in robot bed-making). Depth images are lower dimensional than RGB (one vs three 8-bit channels) and CNNs appear to learn filters for edges and spatial patterns in both.</p>
<h3 id="paper-references">Paper References</h3>
<p>We encourage readers to check out the following papers and project websites for more details.</p>
<ul>
<li>
<p><strong><a href="https://arxiv.org/abs/1703.09312" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics</a>.</strong><br /> Jeffrey Mahler, Jacky Liang, Sherdil Niyaz, Michael Laskey, Richard Doan, Xinyu Liu, Juan Aparicio Ojea, Ken Goldberg.<br /> Robotics: Science and Systems (RSS), 2017.<br /> <a href="https://berkeleyautomation.github.io/dex-net/#dexnet_2" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project Website</a></p>
</li>
<li>
<p><strong><a href="http://proceedings.mlr.press/v78/mahler17a/mahler17a.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Learning Deep Policies for Robot Bin Picking by Simulating Robust Grasping Sequences</a>.</strong><br /> Jeffrey Mahler and Ken Goldberg.<br /> Conference on Robot Learning (CoRL), 2017.<br /> <a href="https://berkeleyautomation.github.io/dex-net/#dexnet_21" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project Website</a></p>
</li>
<li>
<p><strong><a href="https://arxiv.org/abs/1809.05825" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Segmenting Unknown 3D Objects from Real Depth Images using Mask R-CNN Trained on Synthetic Point Clouds</a></strong>.<br /> Michael Danielczuk, Matthew Matl, Saurabh Gupta, Andrew Lee, Andrew Li, Jeffrey Mahler, Ken Goldberg.<br /> arXiv preprint, arXiv:1809.05825<br /> <a href="https://sites.google.com/view/wisdom-dataset/home" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project Website</a></p>
</li>
<li>
<p><strong><a href="https://arxiv.org/abs/1809.09810" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Robot Bed-Making: Deep Transfer Learning Using Depth Sensing of Deformable Fabric</a></strong>.<br /> Daniel Seita*, Nawid Jamali*, Michael Laskey*, Ron Berenstein, Ajay Kumar Tanwani, Prakash Baskaran, Soshi Iba, John Canny, Ken Goldberg.<br /> arXiv preprint, arXiv:1809.09810<br /> <a href="https://sites.google.com/view/bed-make" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Project Website</a></p>
</li>
</ul>
<p>Additional papers and projects can be found at <a href="http://autolab.berkeley.edu/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">the AUTOLab website</a>. This article was initially published on the <a href="https://bair.berkeley.edu/blog/?refresh=1" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Learning acrobatics by watching YouTube</title>
		<link>https://robohub.org/learning-acrobatics-by-watching-youtube/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Thu, 18 Oct 2018 21:11:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/learning-acrobatics-by-watching-youtube/</guid>

					<description><![CDATA[<p>
    <img src="http://bair.berkeley.edu/static/blog/sfv/teaser.gif" alt="..."><br><i>
Simulated characters imitating skills from YouTube videos.
</i>
</p>

<p>Whether it&#8217;s everyday tasks like washing our hands or stunning feats of
acrobatic prowess, humans are able to learn an incredible array of skills by
watching other humans. With the proliferation of publicly available video data
from sources like YouTube, it is now easier than ever to find video clips of
whatever skills we are interested in. A staggering 300 hours of videos are
uploaded to YouTube every minute. Unfortunately, it is still very challenging
for our machines to learn skills from this vast volume of visual data. Most
imitation learning approaches require concise representations, such as those
recorded from motion capture (mocap). But getting mocap data can be quite a
hassle, often requiring heavy instrumentation. Mocap systems also tend to be
restricted to indoor environments with minimal occlusion, which can limit the
types of skills that can be recorded. So wouldn&#8217;t it be nice if our agents can
also learn skills by watching video clips?</p>

<p>In this work, we present a framework for learning skills from videos (SFV). By
combining state-of-the-art techniques in <a href="https://akanazawa.github.io/hmr/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">computer vision</a> and <a href="https://xbpeng.github.io/projects/DeepMimic/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reinforcement
learning</a>, our system enables simulated characters to learn a diverse
repertoire of skills from video clips. Given a single monocular video of an
actor performing some skill, such as a cartwheel or a backflip, our characters
are able to learn policies that reproduce that skill in a physics simulation,
without requiring any manual pose annotations.</p>

<div>
  
</div>

<p><br></p>

<!--more-->

<p>The problem of learning full-body motion skills from videos has received some
attention in computer graphics. Previous <a href="http://graphics.cs.cmu.edu/projects/controllerCapture/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">techniques</a>
often rely on manually-crafted control structures that impose strong
restrictions on the behaviours that can be produced. Therefore, these methods
tend to be limited in the types of skills that can be learned, and the resulting
motions can look fairly unnatural. More recently, deep learning techniques have
demonstrated promising results for visual imitation on domains such as <a href="https://arxiv.org/abs/1805.11592" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Atari</a> and fairly simple <a href="https://arxiv.org/abs/1704.06888" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robotics tasks</a>. But these tasks
often only have modest domain shifts between the demonstrations and the agent&#8217;s
environment, and results on continuous control have largely been on tasks with
relatively simple dynamics.</p>

<h2>Framework</h2>

<p>Our framework is structured as a pipeline, consisting of three stages: pose
estimation, motion reconstruction, and motion imitation. The input video is
first processed by the pose estimation stage, which predicts the pose of the
actor in each frame. Next, the motion reconstruction stage consolidates the pose
predictions into a reference motion and fixes artifacts that might have been
introduced by the pose predictions. Finally, the reference motion is passed to
the motion imitation stage, where a simulated character is trained to imitate
the motion using reinforcement learning.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/sfv/overview.png" width="600" alt="..."><br><i>
The pipeline consists of three stages: pose estimation, motion reconstruction,
and motion imitation. It receives as input, a video clip of an actor performing
a particular skill and a simulated character model, and learns a control policy
that enables the character to reproduce the skill in a physics simulation.
</i>
</p>

<h3>Pose Estimation</h3>

<p>Given a video clip, we use a vision-based pose estimator to predict the actor&#8217;s
pose $\hat{q}_t$ in each frame. The pose estimator is built on the work from <a href="https://akanazawa.github.io/hmr/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">human mesh recovery</a>, which uses a
weakly-supervised adversarial approach to train a pose estimator to predict
poses from monocular images. While pose annotations are required to train the
pose estimator, once trained, the pose estimator can be applied to new images
without any annotations.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/sfv/pose_est_backflip.gif" height="180" alt="..."><img src="http://bair.berkeley.edu/static/blog/sfv/pose_est_handspring.gif" height="180" alt="..."><br><i>
A vision-based pose estimator is used to predict the pose of the actor in each video frame.
</i>
</p>

<h3>Motion Reconstruction</h3>

<p>Since the pose estimator predicts the pose of the actor independently for each
video frame, the predictions between frames can be inconsistent, resulting in
jittery artifacts. Furthermore, while vision-based pose estimators have improved
substantially in recent years, they can still occasionally make some pretty big
mistakes, which can result in peculiar poses popping up every now and then.
These artifacts can produce motions that are physically impossible to imitate.
Therefore, the role of the motion reconstruction stage is to mitigate these
artifacts in order to produce a more physically-plausible reference motion that
will be easier for the simulated character to imitate. To do this, we optimize a
new reference motion  to
satisfy the following objective:</p>

<p>where $l_p(\hat{Q})$ encourages the reference motion to be similar to the
original pose predictions, and $l_{sm}(\hat{Q})$ encourages the poses in
adjacent frames to be similar in order to produce a smoother motion. In
addition, $w_p$ and $w_{sm}$ are the weights for the different losses.</p>

<p>This procedure can substantially improve the quality of the reference motion,
and can fix a lot of the artifacts from the original pose predictions.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/sfv/reconstruction_cartwheel.gif" alt="..."><br><i>
Comparison of reference motions before and after motion reconstruction. Motion
reconstruction mitigates many of the artifacts and produces a smoother reference
motion.
</i>
</p>

<h3>Motion Imitation</h3>

<p>Once we have the reference motion , we can then proceed to training a simulated character to imitate
the skill. The motion imitation stage uses a similar RL approach to the one we
previously proposed for <a href="https://xbpeng.github.io/projects/DeepMimic/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">imitating mocap
data</a>. The reward function simply encourage the policy to minimize the
difference between the pose of the simulated character  and the pose of the
reconstructed reference motion $\hat{q}_t$ at each frame $t$,</p>

<p>Again, this simple approach ends up working surprisingly well, and our
characters are able to learn a diverse repertoire of challenging acrobatic
skills, where each skill is learned from a single video demonstration.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/sfv/humanoid_cartwheel.gif" height="180" alt="..."><img src="http://bair.berkeley.edu/static/blog/sfv/humanoid_frontflip.gif" height="180" alt="..."><br></p>
<p>
    <img src="http://bair.berkeley.edu/static/blog/sfv/humanoid_kipup.gif" height="180" alt="..."><img src="http://bair.berkeley.edu/static/blog/sfv/humanoid_spin.gif" height="180" alt="..."><br><i>
Simulated humanoids learn to perform a diverse array of skills by imitating video clips.
</i>
</p>

<h2>Results</h2>

<p>In total, our characters are able to learn over 20 different skills from various
video clips collected from YouTube.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/sfv/mosiac.gif" alt="..."><br><i>
Our framework can learn a large repertoire of skills from video demonstrations.
</i>
</p>

<p>Even though the morphology of our characters are often quite different from the
actors in the videos, the policies are still able to closely reproduce many of
the skills. As an example of more extreme morphological differences, we can also
train a simulated Atlas robot to imitate video clips of humans.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/sfv/atlas_backflip.gif" height="180" alt="..."><img src="http://bair.berkeley.edu/static/blog/sfv/atlas_handpspring.gif" height="180" alt="..."><br></p>
<p>
    <img src="http://bair.berkeley.edu/static/blog/sfv/atlas_dance.gif" height="180" alt="..."><img src="http://bair.berkeley.edu/static/blog/sfv/atlas_vault.gif" height="180" alt="..."><br><i>
Simulated humanoid learns to perform a diverse array of skills by imitating video clips.
</i>
</p>

<p>One of the advantages of having a simulated character is that we can leverage
the simulation to generalize the behaviours to new environments. Here we have
simulated characters that learn to adapt motions to irregular terrain, where the
original video clips were recorded from actors on flat ground.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/sfv/backflip_slopes.gif" alt="..."><br></p>
<p>
    <img src="http://bair.berkeley.edu/static/blog/sfv/cartwheel_gaps.gif" alt="..."><br><i>
Motions can be adapted to irregular environments.
</i>
</p>

<p>Even though the environments are quite different from those in the original
videos, the learning algorithm still develops fairly plausible strategies for
handling these new environments.</p>

<p>All in all, our framework is really just taking the most obvious approach that
anyone can think of when tackling the problem of video imitation. The key is in
decomposing the problem into more manageable components, picking the right
methods for those components, and integrating them together effectively.
However, imitating skills from videos is still an extremely challenging problem,
and there are plenty of video clips that we are not yet able to reproduce:</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/sfv/gangnam.gif" alt="..."><br><i>
Nimble dance steps, such as this Gangnam style clip, can still be difficult to imitate.
</i>
</p>

<p>But it is encouraging to see that just by integrating together existing
techniques, we can already get pretty far on this challenging problem. We still
have all of our work ahead of us, and we hope that this work will help inspire
future techniques that will enable agents to take advantage of the massive
volume of publicly available video data to acquire a truly staggering array of
skills.</p>

<p>To learn more, <a href="https://xbpeng.github.io/projects/SFV/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">check out our paper and the project webpage</a>.</p>

<p>We would like to thank the co-authors of this work, without whom none of this
would have been possible: Angjoo Kanazawa, Jitendra Malik, Pieter Abbeel, and
Sergey Levine. This research was funded by NSERC, UC Berkeley, BAIR, and AWS.</p>]]></description>
										<content:encoded><![CDATA[<p><strong>By Xue Bin (Jason) Peng and Angjoo Kanazawa</strong></p>
<p>Whether it’s everyday tasks like washing our hands or stunning feats of acrobatic prowess, humans are able to learn an incredible array of skills by watching other humans. With the proliferation of publicly available video data from sources like YouTube, it is now easier than ever to find video clips of whatever skills we are interested in. <span id="more-110197"></span><br />
<img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/teaser.gif" alt="..." />     <br /> <i> Simulated characters imitating skills from YouTube videos. </i> </p>
<p>A staggering 300 hours of videos are uploaded to YouTube every minute. Unfortunately, it is still very challenging for our machines to learn skills from this vast volume of visual data. Most imitation learning approaches require concise representations, such as those recorded from motion capture (mocap). But getting mocap data can be quite a hassle, often requiring heavy instrumentation. Mocap systems also tend to be restricted to indoor environments with minimal occlusion, which can limit the types of skills that can be recorded. So wouldn’t it be nice if our agents can also learn skills by watching video clips?</p>
<p>In this work, we present a framework for learning skills from videos (SFV). By combining state-of-the-art techniques in <a href="https://akanazawa.github.io/hmr/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">computer vision</a> and <a href="https://xbpeng.github.io/projects/DeepMimic/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reinforcement learning</a>, our system enables simulated characters to learn a diverse repertoire of skills from video clips. Given a single monocular video of an actor performing some skill, such as a cartwheel or a backflip, our characters are able to learn policies that reproduce that skill in a physics simulation, without requiring any manual pose annotations.</p>
<div class="keep-aspect"><iframe title="SIGGRAPH Asia 2018: Skills from Videos paper (main video)" width="500" height="281" src="https://www.youtube-nocookie.com/embed/4Qg5I5vhX7Q?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></div>
<p></p>
<p></p>
<p>  <!--more-->  </p>
<p>The problem of learning full-body motion skills from videos has received some attention in computer graphics. Previous <a href="http://graphics.cs.cmu.edu/projects/controllerCapture/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">techniques</a> often rely on manually-crafted control structures that impose strong restrictions on the behaviours that can be produced. Therefore, these methods tend to be limited in the types of skills that can be learned, and the resulting motions can look fairly unnatural. More recently, deep learning techniques have demonstrated promising results for visual imitation on domains such as <a href="https://arxiv.org/abs/1805.11592" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Atari</a> and fairly simple <a href="https://arxiv.org/abs/1704.06888" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robotics tasks</a>. But these tasks often only have modest domain shifts between the demonstrations and the agent’s environment, and results on continuous control have largely been on tasks with relatively simple dynamics.</p>
<h2 id="framework">Framework</h2>
<p>Our framework is structured as a pipeline, consisting of three stages: pose estimation, motion reconstruction, and motion imitation. The input video is first processed by the pose estimation stage, which predicts the pose of the actor in each frame. Next, the motion reconstruction stage consolidates the pose predictions into a reference motion and fixes artifacts that might have been introduced by the pose predictions. Finally, the reference motion is passed to the motion imitation stage, where a simulated character is trained to imitate the motion using reinforcement learning.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/overview.png" width="600" alt="..." />     <br /> <i> The pipeline consists of three stages: pose estimation, motion reconstruction, and motion imitation. It receives as input, a video clip of an actor performing a particular skill and a simulated character model, and learns a control policy that enables the character to reproduce the skill in a physics simulation. </i> </p>
<h3 id="pose-estimation">Pose Estimation</h3>
<p>Given a video clip, we use a vision-based pose estimator to predict the actor’s pose $\hat{q}_t$ in each frame. The pose estimator is built on the work from <a href="https://akanazawa.github.io/hmr/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">human mesh recovery</a>, which uses a weakly-supervised adversarial approach to train a pose estimator to predict poses from monocular images. While pose annotations are required to train the pose estimator, once trained, the pose estimator can be applied to new images without any annotations.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/pose_est_backflip.gif" height="180" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/pose_est_handspring.gif" height="180" alt="..." />     <br /> <i> A vision-based pose estimator is used to predict the pose of the actor in each video frame. </i> </p>
<h3 id="motion-reconstruction">Motion Reconstruction</h3>
<p>Since the pose estimator predicts the pose of the actor independently for each video frame, the predictions between frames can be inconsistent, resulting in jittery artifacts. Furthermore, while vision-based pose estimators have improved substantially in recent years, they can still occasionally make some pretty big mistakes, which can result in peculiar poses popping up every now and then. These artifacts can produce motions that are physically impossible to imitate. Therefore, the role of the motion reconstruction stage is to mitigate these artifacts in order to produce a more physically-plausible reference motion that will be easier for the simulated character to imitate. To do this, we optimize a new reference motion <script type="math/tex">\hat{Q} = \{ \hat{q}_0, \hat{q}_1, \ldots, \hat{q}_t \}</script> to satisfy the following objective:</p>
<p>  <script type="math/tex; mode=display">\min_{\hat{Q}} \; \; w_p l_p(\hat{Q}) + w_{sm} l_{sm}(\hat{Q})</script>  </p>
<p>where $l_p(\hat{Q})$ encourages the reference motion to be similar to the original pose predictions, and $l_{sm}(\hat{Q})$ encourages the poses in adjacent frames to be similar in order to produce a smoother motion. In addition, $w_p$ and $w_{sm}$ are the weights for the different losses.</p>
<p>This procedure can substantially improve the quality of the reference motion, and can fix a lot of the artifacts from the original pose predictions.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/reconstruction_cartwheel.gif" alt="..." />     <br /> <i> Comparison of reference motions before and after motion reconstruction. Motion reconstruction mitigates many of the artifacts and produces a smoother reference motion. </i> </p>
<h3 id="motion-imitation">Motion Imitation</h3>
<p>Once we have the reference motion <script type="math/tex">\{\hat{q}_0, \hat{q}_1, \ldots, \hat{q}_T\}</script>, we can then proceed to training a simulated character to imitate the skill. The motion imitation stage uses a similar RL approach to the one we previously proposed for <a href="https://xbpeng.github.io/projects/DeepMimic/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">imitating mocap data</a>. The reward function simply encourage the policy to minimize the difference between the pose of the simulated character  and the pose of the reconstructed reference motion $\hat{q}_t$ at each frame $t$,</p>
<p>  <script type="math/tex; mode=display">r_t = \exp \Big(-2 \|\hat{q}_t-q_t\|^2 \Big).</script>  </p>
<p>Again, this simple approach ends up working surprisingly well, and our characters are able to learn a diverse repertoire of challenging acrobatic skills, where each skill is learned from a single video demonstration.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/humanoid_cartwheel.gif" height="180" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/humanoid_frontflip.gif" height="180" alt="..." />      </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/humanoid_kipup.gif" height="180" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/humanoid_spin.gif" height="180" alt="..." />     <br /> <i> Simulated humanoids learn to perform a diverse array of skills by imitating video clips. </i> </p>
<h2 id="results">Results</h2>
<p>In total, our characters are able to learn over 20 different skills from various video clips collected from YouTube.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/mosiac.gif" alt="..." />     <br /> <i> Our framework can learn a large repertoire of skills from video demonstrations. </i> </p>
<p>Even though the morphology of our characters are often quite different from the actors in the videos, the policies are still able to closely reproduce many of the skills. As an example of more extreme morphological differences, we can also train a simulated Atlas robot to imitate video clips of humans.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/atlas_backflip.gif" height="180" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/atlas_handpspring.gif" height="180" alt="..." />      </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/atlas_dance.gif" height="180" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/atlas_vault.gif" height="180" alt="..." />     <br /> <i> Simulated humanoid learns to perform a diverse array of skills by imitating video clips. </i> </p>
<p>One of the advantages of having a simulated character is that we can leverage the simulation to generalize the behaviours to new environments. Here we have simulated characters that learn to adapt motions to irregular terrain, where the original video clips were recorded from actors on flat ground.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/backflip_slopes.gif" alt="..." />      </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/cartwheel_gaps.gif" alt="..." />     <br /> <i> Motions can be adapted to irregular environments. </i> </p>
<p>Even though the environments are quite different from those in the original videos, the learning algorithm still develops fairly plausible strategies for handling these new environments.</p>
<p>All in all, our framework is really just taking the most obvious approach that anyone can think of when tackling the problem of video imitation. The key is in decomposing the problem into more manageable components, picking the right methods for those components, and integrating them together effectively. However, imitating skills from videos is still an extremely challenging problem, and there are plenty of video clips that we are not yet able to reproduce:</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/sfv/gangnam.gif" alt="..." />     <br /> <i> Nimble dance steps, such as this Gangnam style clip, can still be difficult to imitate. </i> </p>
<p>But it is encouraging to see that just by integrating together existing techniques, we can already get pretty far on this challenging problem. We still have all of our work ahead of us, and we hope that this work will help inspire future techniques that will enable agents to take advantage of the massive volume of publicly available video data to acquire a truly staggering array of skills.</p>
<p>To learn more, <a href="https://xbpeng.github.io/projects/SFV/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">check out our paper and the project webpage</a>.</p>
<p>We would like to thank the co-authors of this work, without whom none of this would have been possible: Angjoo Kanazawa, Jitendra Malik, Pieter Abbeel, and Sergey Levine. This research was funded by NSERC, UC Berkeley, BAIR, and AWS. This article was initially published on the <a href="https://bair.berkeley.edu/blog" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Dexterous manipulation with reinforcement learning: Efficient, general, and low-cost</title>
		<link>https://robohub.org/dexterous-manipulation-with-reinforcement-learning-efficient-general-and-low-cost/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Mon, 10 Sep 2018 11:07:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/dexterous-manipulation-with-reinforcement-learning-efficient-general-and-low-cost/</guid>

					<description><![CDATA[    
    


In this post, we demonstrate how deep reinforcement learning (deep RL) can be
used to learn how to control dexterous hands for a variety of manipulation
tasks. We discuss how such methods can learn to make use of low-cost hardware,
can be ...]]></description>
										<content:encoded><![CDATA[<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/teaser.gif" height="300" alt="..." />      </p>
<p>In this post, we demonstrate how deep reinforcement learning (deep RL) can be used to learn how to control dexterous hands for a variety of manipulation tasks. We discuss how such methods can learn to make use of low-cost hardware, can be implemented efficiently, and how they can be complemented with techniques such as demonstrations and simulation to accelerate learning.</p>
<p>  <span id="more-107481"></span>  </p>
<h2 id="why-dexterous-hands">Why Dexterous Hands?</h2>
<p>A majority of robots in use today use simple parallel jaw grippers as manipulators, which are sufficient for structured settings like factories. However, manipulators that are capable of performing a wide array of tasks are essential for unstructured human-centric environments like the home. Multi-fingered hands are among the most versatile manipulators, and enable a wide variety of skills we use in our everyday life such as moving objects, opening doors, typing, and painting.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/missing_img2.png" height="180" alt="..." />      </p>
<p>Unfortunately, controlling dexterous hands is extremely difficult, which limits their use. High-end hands can also be extremely expensive, due to delicate sensing and actuation. Deep reinforcement learning offers the promise of automating complex control tasks even with cheap hardware, but many applications of deep RL use huge amounts of simulated data, making them expensive to deploy in terms of both cost and engineering effort.  Humans can learn motor skills efficiently, without a simulator and millions of simulations.</p>
<p>We will first show that deep RL can in fact be used to learn complex manipulation behaviors by training directly in the real world, with modest computation and low-cost robotic hardware, and without any model or simulator. We then describe how learning can be further accelerated by incorporating additional sources of supervision, including demonstrations and simulation. We demonstrate learning on two separate hardware platforms: an inexpensive custom-built 3-fingered hand (the Dynamixel Claw), which costs under \$2500, and the higher-end Allegro hand, which costs about \$15,000.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/missing_img1.png" height="230" alt="..." />     <br />     <i>     Left: Dynamixel Claw. Right: Allegro Hand.     </i> </p>
<h2 id="model-free-reinforcement-learning-in-the-real-world">Model-free Reinforcement Learning in the Real World</h2>
<p>Deep RL algorithms learn by trial and error, maximizing a user-specified reward function from experience. We’ll use a valve rotation task as a working example, where the hand must open a valve or faucet by rotating it 180 degrees.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/hand_valve_rotation.gif" height="300" alt="..." />     <br />     <i>     Illustration of valve rotation task.     </i> </p>
<p>The reward function simply consists of the negative distance between the current and desired valve orientation, and the hand must figure out on its own how to move to rotate it. A central challenge in deep RL is in using this weak reward signal to find a complex and coordinated behavior strategy (a <em>policy</em>) that succeeds at the task. The policy is represented by a multilayer neural network. This typically requires a large number of trials, so much so that the community is split on whether deep RL methods can be used for training outside of simulation.  However, this imposes major limitations on their applicability: learning directly in the real world makes it possible to learn any task from experience, while using simulators requires designing a suitable simulation, modeling the task and the robot, and carefully adjusting their parameters to achieve good results. We will show later that simulation can accelerate learning substantially, but we first demonstrate that existing RL algorithms can in fact learn this task directly on real hardware.</p>
<p>A variety of algorithms should be suitable. We use <a href="https://arxiv.org/abs/1703.02660" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Truncated Natural Policy Gradient</a> to learn the task, which requires about 9 hours on real hardware.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/iter40_valve.gif" height="230" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/iter80_valve.gif" height="230" alt="..." />     <br />     <i>     Learning progress of the dynamixel claw on valve rotation.     </i> </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/real_rl_valve.gif" height="300" alt="..." />     <br />     <i>     Final Trained Policy on valve rotation.     </i> </p>
<p>The direct RL approach is appealing for a number of reasons. It requires minimal assumptions, and is thus well suited to autonomously acquire a large repertoire of skills. Since this approach assumes no information other than access to a reward function, it is easy to relearn the skill in a modified environment, for example when using a different object or a different hand – in this case, the Allegro hand.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/allegrohand.gif" height="420" alt="..." />     <br />     <i>     360° valve rotation with Allegro Hand.     </i> </p>
<p>The same exact method can learn to rotate the valve when we use a different material. We can learn how to rotate a valve made out of foam. This can be quite difficult to simulate accurately, and training directly in the real world allows us to learn without needing accurate simulations.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/foamscrew.gif" height="360" alt="..." />     <br />     <i>     Dynamixel claw rotating a foam screw.     </i> </p>
<p>The same approach takes 8 hours to solve a different task, which requires flipping an object 180 degrees around the horizontal axis, without any modification.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/flip_hand.gif" height="230" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/real_rl_flipping.gif" height="230" alt="..." />     <br />     <i>     Dynamixel claw flipping a block.     </i> </p>
<p>These behaviors were learned with low cost hardware (&lt;\$2500) and a single consumer desktop computer.</p>
<h2 id="accelerating-learning-with-human-demonstrations">Accelerating Learning with Human Demonstrations</h2>
<p>While model-free RL is extremely general, incorporating supervision from human experts can help accelerate learning further. One way to do this, which we describe in our paper on <a href="https://arxiv.org/abs/1709.10087" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Demonstration Augmented Policy Gradient (DAPG)</a>, is to incorporate human demonstrations into the reinforcement learning process. Related approaches have been proposed in the context of <a href="https://arxiv.org/abs/1709.10089" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">off-policy RL</a>, <a href="https://arxiv.org/abs/1704.03732" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Q-learning</a>, and <a href="http://is.tuebingen.mpg.de/fileadmin/user_upload/files/publications/ICRA2009-Kober_5661%5B0%5D.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">other robotic tasks</a>. The key idea behind DAPG is that demonstrations can be used to accelerate RL in two ways</p>
<img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/missing_img3.png" height="300" align="right" hspace="30" alt="..." />
<ol>
<li>
<p>Provide a good initialization for the policy via behavior cloning.</p>
</li>
<li>
<p>Provide an auxiliary learning signal <em>throughout</em> the learning process to guide exploration using a trajectory tracking auxiliary reward.</p>
</li>
</ol>
<p>The auxiliary objective during RL prevents the policy from diverging from the demonstrations during the RL process. Pure behavior cloning with limited data is often ineffective in training successful policies due to distribution drift and limited data support. RL is crucial for robustness and generalization and use of demonstrations can substantially accelerate the learning process.  We previously validated this algorithm in simulation on a variety of tasks, shown below, where each task used only 25 human demonstrations collected in virtual reality.  DAPG enables a speedup of up to 30x on these tasks, while also learning natural and robust behaviors.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/task_relocate.gif" height="160" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/task_hammer.gif" height="160" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/task_pen.gif" height="160" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/task_door.gif" height="160" alt="..." />     <br />     <i>     Behaviors learned in simulation with DAPG: object pickup, tool use, in-hand,     door opening.     </i> </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/dapg_robustness.gif" height="230" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/pure_rl_robustness.gif" height="230" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/natural_motion_dapg.gif" height="230" alt="..." />     <br />     <i>     Behaviors robust to size and shape variations; natural and smooth behavior.     </i> </p>
<p>In the real world, we can use this algorithm with the dynamixel claw to significantly accelerate learning. The demonstrations are collected with kinesthetic teaching, where a human teacher moves the fingers of the robots directly in the real world. This brings down the training time on both tasks to under 4 hours.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/dapg_valve.gif" height="230" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/dapg_flipping.gif" height="230" alt="..." />     <br />     <i>     Left: Valve rotation policy with DAPG. Right: Flipping policy with DAPG.     </i> </p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/missing_img4.png" height="320" alt="..." />     <br />     <i>     Learning Curves of RL from scratch on hardware vs DAPG.     </i> </p>
<p>Demonstrations provide a natural way to incorporate human priors and accelerate the learning process. Where high quality successful demonstrations are available, augmenting RL with demonstrations has the potential to substantially accelerate RL. However, obtaining demonstrations may not be possible for all tasks or robot morphologies, necessitating the need to also pursue alternate acceleration schemes.</p>
<h2 id="accelerating-learning-with-simulation">Accelerating Learning with Simulation</h2>
<p>A simulated model of the task can help augment the real world data with large amounts of simulated data to accelerate the learning process. For the simulated data to be representative of the complexities of the real world, randomization of various simulation parameters is often necessitated. This kind of randomization has previously been observed to produce <a href="https://arxiv.org/abs/1610.01283" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robust</a> policies, and can facilitate transfer in the face of both <a href="https://arxiv.org/abs/1611.04201" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">visual</a> and <a href="https://xbpeng.github.io/projects/SimToReal/2018_SimToReal.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">physical</a> <a href="https://arxiv.org/abs/1803.10371" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">discrepancies</a>.  Our experiments also suggest that simulation to reality transfer with randomization can be effective.</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/sim2real.gif" height="320" alt="..." />     <br />     <i>     Policy for valve rotation transferred from simulation using randomization.     </i> </p>
<p>Transfer from simulation has also been explored in <a href="https://arxiv.org/abs/1808.00177" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">concurrent work</a> for dexterous manipulation to learn impressive behaviors, and in a number of prior works for tasks such as <a href="https://arxiv.org/abs/1707.02267" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">picking and placing</a>, <a href="https://arxiv.org/abs/1712.07642" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">visual servoing</a>, and <a href="https://arxiv.org/abs/1804.10332" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">locomotion</a>. While simulation to real transfer enabled by randomization is an appealing option, especially for fragile robots, it has a number of limitations. First, the resulting policies can end up being overly conservative due to the randomization, a phenomenon that has been widely observed in the field of robust control. Second, the particular choice of parameters to randomize is crucial for good results, and insights from one task or problem domain may not transfer to others. Third, increasing the amount of randomization results in more complex models tremendously increasing the training time and required computational resources (<a href="https://arxiv.org/abs/1808.00177" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Andrychowicz et al</a> report 100 years of simulated experience, training in 50 hours on thousands of CPU cores). Directly training in the real world may be more efficient and lead to better policies. Finally, and perhaps most importantly, an accurate simulator must be constructed manually, with each new task modeled by hand in the simulation, which requires substantial time and expertise. However, leveraging simulations appropriately can accelerate the learning, and more systematic transfer methods are an important direction for future work.</p>
<h2 id="accelerating-learning-with-learned-models">Accelerating Learning with Learned Models</h2>
<p>In some of <a href="https://homes.cs.washington.edu/~todorov/papers/KumarICRA16.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our previous work</a>, we also studied how learned dynamics models can accelerate real-world reinforcement learning without access to manually engineered simulators. In this approach, local derivatives of the dynamics are approximated by fitting time-varying linear systems, which can then be used to locally and iterative improve a policy. This approach can acquire a variety of in-hand manipulation strategies from scratch in the real world. Furthermore, we see that the same algorithm can even <a href="https://arxiv.org/abs/1603.06348" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">learn to control</a> a pneumatic soft robotic hand to perform a number of dexterous behaviors</p>
<p style="text-align:center;">     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/learned_local_models_adroit.gif" height="230" alt="..." />     <img decoding="async" src="http://bair.berkeley.edu/static/blog/dex_manip/softhand.gif" height="230" alt="..." />     <br />     <i>     Left: Adroit robotic hand performing in-hand manipulation. Right: Pneumatic     Soft RBO Hand performing dexterous tasks.     </i> </p>
<p>However, the performance of methods with learned models is limited by the quality of the model that can be learned, and in practice asymptotic performance is often still higher with the best model-free algorithms. Further study of model-based reinforcement learning for efficient and effective real-world learning is a promising research direction.</p>
<h2 id="takeaways-and-challenges">Takeaways and Challenges</h2>
<p>While training in the real world is general and broadly applicable, it has several challenges of its own.</p>
<ol>
<li>
<p>Due to the requirement to take a large number of exploratory actions, we observed that the hands often heat up quickly, which requires pauses to avoid damage.</p>
</li>
<li>
<p>Since the hands must attempt the task multiple times, we had to build an automatic reset mechanism. In the future, a promising direction to remove this requirement is to <a href="https://arxiv.org/abs/1711.06782" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">automatically learn</a> <a href="http://rll.berkeley.edu/reset_controller/reset_controller.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reset policies</a>.</p>
</li>
<li>
<p>Reinforcement learning methods require rewards to be provided, and this reward must still be designed manually. Some of our recent work has looked at automating reward specification.</p>
</li>
</ol>
<p>However, enabling robots to learn complex skills directly in the real world is one of the best paths forward to developing truly generalist robots. In the same way that humans can learn directly from experience in the real world, robots that can acquire skills simply by trial and error can explore novel solutions to difficult manipulation problems and discover them with minimal human intervention. At the same time, the availability of demonstrations, simulators, and other prior knowledge can further reduce training times.</p>
<hr />
<p>This article was initially published on the BAIR blog, and appears here with the authors’ permission. The work in this post is based on these papers:</p>
<ul>
<li><a href="https://homes.cs.washington.edu/~todorov/papers/KumarICRA16.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Optimal control with learned local models: Application to dexterous manipulation</a></li>
<li><a href="https://arxiv.org/abs/1603.06348" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Learning Dexterous Manipulation for a Soft Robotic Hand from Human Demonstration</a></li>
<li><a href="https://arxiv.org/abs/1709.10087" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations</a></li>
</ul>
<p><em>A complete paper on the new robotic experiments will be released soon. The research was conducted by Henry Zhu, Abhishek Gupta, Vikash Kumar, Aravind Rajeswaran, and Sergey Levine. Collaborators on earlier projects include Emo Todorov, John Schulman, Giulia Vezzani, Pieter Abbeel, Clemens Eppner.</em></p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>When recurrent models don&#8217;t need to be recurrent</title>
		<link>https://robohub.org/when-recurrent-models-dont-need-to-be-recurrent/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Thu, 09 Aug 2018 13:59:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">https://robohub.org/when-recurrent-models-dont-need-to-be-recurrent/</guid>

					<description><![CDATA[<p><em>An earlier version of this post was published on <a href="http://www.offconvex.org/2018/07/27/approximating-recurrent/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Off the Convex
Path</a>. It is reposted here with the
author&#8217;s permission.</em></p>

<p>In the last few years, deep learning practitioners have proposed a litany of
different sequence models.  Although recurrent neural networks were once the
tool of choice, now models like the autoregressive
<a href="https://deepmind.com/blog/wavenet-generative-model-raw-audio/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wavenet</a> or the
<a href="https://ai.googleblog.com/2017/08/transformer-novel-neural-network.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Transformer</a>
are replacing RNNs on a diverse set of tasks. In this post, we explore the
trade-offs between recurrent and feed-forward models. Feed-forward models can
offer improvements in training stability and speed, while recurrent models are
strictly more expressive. Intriguingly, this added expressivity does not seem to
boost the performance of recurrent models.  Several groups have shown
feed-forward networks can match the results of the best recurrent models on
benchmark sequence tasks. This phenomenon raises an interesting question for
theoretical investigation:</p>

<blockquote>
  <p>When and why can feed-forward networks replace recurrent neural networks
without a loss in performance?</p>
</blockquote>

<p>We discuss several proposed answers to this question and highlight our
<a href="https://arxiv.org/abs/1805.10369" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recent work</a> that offers an explanation in
terms of a fundamental stability property.</p>

<!--more-->

<h1>A Tale of Two Sequence Models</h1>
<h2>Recurrent Neural Networks</h2>
<p>The many variants of recurrent models all have a similar form. The model
maintains a state $h_t$ that summarizes the past sequence of inputs. At each
time step $t$, the state is updated according to the equation</p>

<p>where $x_t$ is the input at time $t$, $\phi$ is a differentiable map, and $h_0$
is an initial state. In a vanilla recurrent neural network, the model is
parameterized by matrices $W$ and $U$, and the state is updated according to</p>

<p>In practice, the <a href="http://colah.github.io/posts/2015-08-Understanding-LSTMs/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Long Short-Term Memory
(LSTM)</a> network is
more frequently used. In either case, to make predictions, the state is passed
to a function $f$, and the model predicts $y_t = f(h_t)$. Since the state $h_t$
is a function of all of the past inputs $x_0, \dots, x_t$, the prediction $y_t$
depends on the entire history $x_0, \dots, x_t$ as well.</p>

<p>A recurrent model can also be represented graphically.</p>

<p>
    <img src="http://bair.berkeley.edu/static/blog/recurrent/recurrent_net.png" width="500px" height="250px"></p>

<p>Recurrent models are fit to data using backpropagation. However, backpropagating
gradients from time step $T$ to time step $0$ often requires infeasibly large
amounts of memory, so essentially every implementation of a recurrent model
<em>truncates</em> the model and only backpropagates gradient $k$ times steps.</p>
<figure><p>
        <img src="http://bair.berkeley.edu/static/blog/recurrent/truncated_backprop.png"></p>
    <figcaption><small>
        Source: <a href="https://r2rt.com/styles-of-truncated-backpropagation.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"> 
        https://r2rt.com/styles-of-truncated-backpropagation.html </a>
    </small>
    </figcaption></figure><p>In this setup, the predictions of the recurrent model still depend on the entire
history $x_0, \dots, x_T$. However, it&#8217;s not clear how this training procedure
affects the model&#8217;s ability to learn long-term patterns, particularly those that
require more than $k$ steps.</p>

<h2>Autoregressive, Feed-Forward Models</h2>
<p>Instead of making predictions from a state that depends on the entire history,
an autoregressive model directly predicts $y_t$ using only the $k$ most recent
inputs, $x_{t-k+1}, \dots, x_{t}$. This corresponds to a strong <em>conditional
independence</em> assumption. In particular, a feed-forward model assumes the target
only depends on the $k$ most recent inputs. Google&#8217;s
<a href="https://arxiv.org/abs/1609.03499" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">WaveNet</a> nicely illustrates this general
principle.</p>

<figure><p>
        <img src="https://storage.googleapis.com/deepmind-live-cms/documents/BlogPost-Fig2-Anim-160908-r01.gif"></p>
    <figcaption><small>
        Source: <a href="https://deepmind.com/blog/wavenet-generative-model-raw-audio/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"> 
        https://deepmind.com/blog/wavenet-generative-model-raw-audio/</a>
    </small>
    </figcaption></figure><p>In contrast to an RNN, the limited context of a feed-forward model means that it
cannot capture patterns that extend more than $k$ steps. However, using
techniques like dilated-convolutions, one can make $k$ quite large.</p>

<h1>Why Care About Feed-Forward Models?</h1>
<p>At the outset, recurrent models appear to be a strictly more flexible and
expressive model class than feed-forward models. After all, feed-forward
networks make a strong conditional independence assumption that recurrent models
don&#8217;t make. Even if feed-forward models are less expressive, there are still
several reasons one might prefer a feed-forward network.</p>
<ul><li><strong>Parallelization</strong>: Convolutional feed-forward models are easier to <a href="https://arxiv.org/abs/1705.03122" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">parallelize
at training time</a>. 
There&#8217;s no hidden state to update and maintain, and
therefore no sequential dependencies between outputs. This allows very
efficient implementations of training on modern hardware.</li>
  <li><strong>Trainability</strong>: Training deep convolutional neural networks is the
bread-and-butter of deep learning. Whereas recurrent models are often more
finicky and difficult to <a href="https://arxiv.org/abs/1211.5063" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">optimize</a>,
significant effort has gone into designing architectures and software to
efficiently and reliably train deep feed-forward networks.</li>
  <li><strong>Inference Speed</strong>: In some cases, feed-forward models can be significantly
more light-weight and perform <a href="https://arxiv.org/abs/1211.5063" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">inference faster than similar recurrent
systems</a>. In other cases,
particularly for long sequences, autoregressive inference is a large
bottleneck and requires <a href="https://arxiv.org/abs/1702.07825" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">significant engineering
work</a> or <a href="https://arxiv.org/abs/1711.10433" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">significant
cleverness</a> to overcome.</li>
</ul><h1>Feed-Forward Models Can Outperform Recurrent Models</h1>
<p>Although it appears trainability and parallelization for feed-forward models
comes at the price of reduced accuracy, there have been several recent examples
showing that feed-forward networks can actually achieve the same accuracies as
their recurrent counterparts on benchmark tasks.</p>

<ul><li>
    <p><strong>Language Modeling.</strong>
In language modeling, the goal is to predict the next word in a document given
all of the previous words. Feed-forward models make predictions using only the
$k$ most recent words, whereas recurrent models can potentially use the entire
document.  The <a href="https://arxiv.org/abs/1612.08083" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gated-Convolutional Language
Model</a> is a feed-forward autoregressive models
that is competitive with <a href="https://arxiv.org/abs/1602.02410" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">large LSTM baseline
models</a>. Despite using a truncation length of
$k=25$, the model outperforms a large LSTM on the
<a href="https://einstein.ai/research/the-wikitext-long-term-dependency-language-modeling-dataset" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wikitext-103</a>
benchmark, which is designed to reward models that capture long-term
dependencies. On the <a href="http://www.statmt.org/lm-benchmark/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Billion Word
Benchmark</a>, the model is slightly worse
than the largest LSTM, but is faster to train and uses fewer resources.</p>
  </li>
  <li>
    <p><strong>Machine Translation.</strong>
The goal in machine translation is to map sequences of English words to
sequences of, say, French words. Feed-forward models make translations using
only $k$ words of the sentence, whereas recurrent models can leverage the entire
sentence.  Within the deep learning world, variants of the LSTM-based <a href="https://arxiv.org/abs/1409.0473" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sequence
to Sequence with Attention</a> model, particularly
<a href="https://arxiv.org/abs/1609.08144" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Google Neural Machine Translation</a>, were
superseded first by a fully <a href="https://arxiv.org/abs/1705.03122" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">convolutional sequence to
sequence</a> model and then by the
<a href="https://arxiv.org/abs/1706.03762" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Transformer</a>.<sup><a href="http://bair.berkeley.edu/blog/#fn:1" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">1</a></sup></p>
  </li>
</ul><figure><p>
        <img src="https://raw.githubusercontent.com/facebookresearch/fairseq/master/fairseq.gif"></p>
    <figcaption><small>
        Source: <a href="https://github.com/facebookresearch/fairseq/blob/master/fairseq.gif" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"> 
        https://github.com/facebookresearch/fairseq/blob/master/fairseq.gif </a>
    </small>
    </figcaption></figure><ul><li>
    <p><strong>Speech Synthesis.</strong>
In speech synthesis, one seeks to generate a realistic human speech signal.
Feed-forward models are limited to the past $k$ samples, whereas recurrent
models can use the entire history. Upon publication, the feed-forward,
autoregressive <a href="https://arxiv.org/abs/1609.03499" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">WaveNet</a> was a substantial
improvement over LSTM-RNN parametric models.</p>
  </li>
  <li>
    <p><strong>Everthing Else.</strong> 
Recently <a href="https://arxiv.org/abs/1803.01271" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Bai et al.</a> proposed a generic
feed-forward model leveraging dilated convolutions and showed it outperforms
recurrent baselines on tasks ranging from synthetic copying tasks to music
generation.</p>
  </li>
</ul><h1>How Can Feed-Forward Models Outperform Recurrent Ones?</h1>
<p>In the examples above, feed-forward networks achieve results on par with or
better than recurrent networks. This is perplexing since recurrent models
seem to be more powerful a priori. One explanation for this phenomenon is
given by <a href="https://arxiv.org/abs/1612.08083" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Dauphin et al.</a>:</p>

<blockquote>
  <p>The unlimited context offered by recurrent models is not strictly necessary
for language modeling.</p>
</blockquote>

<p>In other words, it&#8217;s possible you don&#8217;t need a large amount of context to do
well on the prediction task on average. <a href="https://arxiv.org/abs/1612.02526" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Recent theoretical
work</a> offers some evidence in favor of this view.</p>

<p>Another explanation is given by <a href="https://arxiv.org/abs/1803.01271" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Bai et al.</a>:</p>
<blockquote>
  <p>The &#8220;infinite memory&#8221; advantage of RNNs is largely absent in practice.</p>
</blockquote>

<p>As Bai et al. report, even in experiments explicitly requiring long-term
context, RNN variants were unable to learn long sequences. On the Billion Word
Benchmark, an <a href="https://arxiv.org/abs/1703.10724" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">intriguing Google Technical
Report</a> suggests an LSTM $n$-gram model with
$n=13$ words of memory is as good as an LSTM with arbitrary context.</p>

<p>This evidence leads us to conjecture: <strong>Recurrent models <em>trained in practice</em>
are effectively feed-forward.</strong> This could happen either because truncated
backpropagation through time cannot learn patterns significantly longer than $k$
steps, or, more provocatively, because models <em>trainable by gradient descent</em>
cannot have long-term memory.</p>

<p>In <a href="https://arxiv.org/abs/1805.10369" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our recent paper</a>, we study the gap
between recurrent and feed-forward models trained using gradient descent. We
show if the recurrent model is <em>stable</em> (meaning the gradients can not explode),
then the model can be well-approximated by a feed-forward network for the
purposes of both <em>inference and training.</em> In other words, we show feed-forward
and stable recurrent models trained by gradient descent are <em>equivalent</em> in the
sense of making identical predictions at test-time.</p>

<p>Stability is a natural criterion for learnability of recurrent models. Outside
of the stable regime, gradient descent cannot be expected to work. Indeed, even
for very simple unstable models, gradient descent fails to converge to a
stationary point. While models trained in practice are not necessarily stable,
the performance of unstable models is likely in spite of, not due to, their
instability.</p>

<p>Using the <a href="https://einstein.ai/research/the-wikitext-long-term-dependency-language-modeling-dataset" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wikitext-2 language modeling
benchmark</a>,
we show stability can be imposed on benchmark models without a loss in
performance. Concretely, we conducted a hyperparameter search to find the
best-performing unstable RNN and LSTM. Then, we re-trained both models while
enforcing the stability conditions derived in our paper. <strong>In both cases, the
unstable and stable models have the same test performance!</strong></p>

<table><thead><tr><th>Recurrent Model</th>
      <th>Unstable (perplexity)</th>
      <th>Stable  (perplexity)</th>
    </tr></thead><tbody><tr><td>Tanh-RNN</td>
      <td>146.7</td>
      <td>143.5</td>
    </tr><tr><td>LSTM</td>
      <td>92.3</td>
      <td>95.1</td>
    </tr></tbody></table><h1>Conclusion</h1>
<p>Despite some initial attempts, there is still much to do to understand
why feed-forward models are competitive with recurrent ones and
shed light onto the trade-offs between sequence models. How much memory is
really needed to perform well on common sequence benchmarks? What are the
expressivity trade-offs between truncated RNNs (which can be considered
feed-forward) and the widely-used convolutional models?</p>

<p>Answering these questions is a step towards building a theory that can both
explain the strengths and limitations of our current methods and give guidance
about how to choose between different classes of models in concrete settings.</p>

<div>
  <ol><li>
      <p>The Transformer isn&#8217;t strictly a feed-forward model in the style described above (since it doesn&#8217;t make the $k$ step conditional independence assumption), but is not really a recurrent model because it doesn&#8217;t maintain a hidden state.&#160;<a href="http://bair.berkeley.edu/blog/#fnref:1" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">&#8617;</a></p>
    </li>
  </ol></div>]]></description>
										<content:encoded><![CDATA[<p><strong>By John Miller</strong></p>
<p><em>An earlier version of this post was published on <a href="http://www.offconvex.org/2018/07/27/approximating-recurrent/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Off the Convex  Path</a>. It is reposted here with the  author’s permission.</em></p>
<p>In the last few years, deep learning practitioners have proposed a litany of  different sequence models.  Although recurrent neural networks were once the  tool of choice, now models like the autoregressive  <a href="https://deepmind.com/blog/wavenet-generative-model-raw-audio/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wavenet</a> or the  <a href="https://ai.googleblog.com/2017/08/transformer-novel-neural-network.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Transformer</a>  are replacing RNNs on a diverse set of tasks. In this post, we explore the  trade-offs between recurrent and feed-forward models.<span id="more-106024"></span></p>
<p> Feed-forward models can  offer improvements in training stability and speed, while recurrent models are  strictly more expressive. Intriguingly, this added expressivity does not seem to  boost the performance of recurrent models.  Several groups have shown  feed-forward networks can match the results of the best recurrent models on  benchmark sequence tasks. This phenomenon raises an interesting question for  theoretical investigation:</p>
<blockquote>
<p>When and why can feed-forward networks replace recurrent neural networks  without a loss in performance?</p>
</blockquote>
<p>We discuss several proposed answers to this question and highlight our  <a href="https://arxiv.org/abs/1805.10369" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">recent work</a> that offers an explanation in  terms of a fundamental stability property.</p>
<p>    <!--more-->    </p>
<h1 id="a-tale-of-two-sequence-models">A Tale of Two Sequence Models</h1>
<h2 id="recurrent-neural-networks">Recurrent Neural Networks</h2>
<p>The many variants of recurrent models all have a similar form. The model  maintains a state $h_t$ that summarizes the past sequence of inputs. At each  time step $t$, the state is updated according to the equation</p>
<p>    <script type="math/tex; mode=display">h_{t+1} = \phi(h_t, x_t),</script>    </p>
<p>where $x_t$ is the input at time $t$, $\phi$ is a differentiable map, and $h_0$  is an initial state. In a vanilla recurrent neural network, the model is  parameterized by matrices $W$ and $U$, and the state is updated according to</p>
<p>    <script type="math/tex; mode=display">h_{t+1} = \tanh(Wh_t + Ux_t).</script>    </p>
<p>In practice, the <a href="http://colah.github.io/posts/2015-08-Understanding-LSTMs/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Long Short-Term Memory  (LSTM)</a> network is  more frequently used. In either case, to make predictions, the state is passed  to a function $f$, and the model predicts $y_t = f(h_t)$. Since the state $h_t$  is a function of all of the past inputs $x_0, \dots, x_t$, the prediction $y_t$  depends on the entire history $x_0, \dots, x_t$ as well.</p>
<p>A recurrent model can also be represented graphically.</p>
<p style="text-align:center;">      <img decoding="async" src="http://bair.berkeley.edu/static/blog/recurrent/recurrent_net.png" width="500px" height="250px" />  </p>
<p>Recurrent models are fit to data using backpropagation. However, backpropagating  gradients from time step $T$ to time step $0$ often requires infeasibly large  amounts of memory, so essentially every implementation of a recurrent model  <em>truncates</em> the model and only backpropagates gradient $k$ times steps.</p>
<figure>
<p style="text-align:center;">          <img decoding="async" src="http://bair.berkeley.edu/static/blog/recurrent/truncated_backprop.png" />      </p><figcaption>      <small>          Source: <a href="https://r2rt.com/styles-of-truncated-backpropagation.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">           https://r2rt.com/styles-of-truncated-backpropagation.html </a>      </small>      </figcaption></figure>
<p>In this setup, the predictions of the recurrent model still depend on the entire  history $x_0, \dots, x_T$. However, it’s not clear how this training procedure  affects the model’s ability to learn long-term patterns, particularly those that  require more than $k$ steps.</p>
<h2 id="autoregressive-feed-forward-models">Autoregressive, Feed-Forward Models</h2>
<p>Instead of making predictions from a state that depends on the entire history,  an autoregressive model directly predicts $y_t$ using only the $k$ most recent  inputs, $x_{t-k+1}, \dots, x_{t}$. This corresponds to a strong <em>conditional  independence</em> assumption. In particular, a feed-forward model assumes the target  only depends on the $k$ most recent inputs. Google’s  <a href="https://arxiv.org/abs/1609.03499" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">WaveNet</a> nicely illustrates this general  principle.</p>
<figure>
<p style="text-align:center;">          <img decoding="async" src="https://storage.googleapis.com/deepmind-live-cms/documents/BlogPost-Fig2-Anim-160908-r01.gif" />      </p><figcaption>      <small>          Source: <a href="https://deepmind.com/blog/wavenet-generative-model-raw-audio/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">           https://deepmind.com/blog/wavenet-generative-model-raw-audio/</a>      </small>      </figcaption></figure>
<p>In contrast to an RNN, the limited context of a feed-forward model means that it  cannot capture patterns that extend more than $k$ steps. However, using  techniques like dilated-convolutions, one can make $k$ quite large.</p>
<h1 id="why-care-about-feed-forward-models">Why Care About Feed-Forward Models?</h1>
<p>At the outset, recurrent models appear to be a strictly more flexible and  expressive model class than feed-forward models. After all, feed-forward  networks make a strong conditional independence assumption that recurrent models  don’t make. Even if feed-forward models are less expressive, there are still  several reasons one might prefer a feed-forward network.</p>
<ul>
<li><strong>Parallelization</strong>: Convolutional feed-forward models are easier to <a href="https://arxiv.org/abs/1705.03122" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">parallelize  at training time</a>.   There’s no hidden state to update and maintain, and  therefore no sequential dependencies between outputs. This allows very  efficient implementations of training on modern hardware.</li>
<li><strong>Trainability</strong>: Training deep convolutional neural networks is the  bread-and-butter of deep learning. Whereas recurrent models are often more  finicky and difficult to <a href="https://arxiv.org/abs/1211.5063" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">optimize</a>,  significant effort has gone into designing architectures and software to  efficiently and reliably train deep feed-forward networks.</li>
<li><strong>Inference Speed</strong>: In some cases, feed-forward models can be significantly  more light-weight and perform <a href="https://arxiv.org/abs/1211.5063" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">inference faster than similar recurrent  systems</a>. In other cases,  particularly for long sequences, autoregressive inference is a large  bottleneck and requires <a href="https://arxiv.org/abs/1702.07825" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">significant engineering  work</a> or <a href="https://arxiv.org/abs/1711.10433" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">significant  cleverness</a> to overcome.</li>
</ul>
<h1 id="feed-forward-models-can-outperform-recurrent-models">Feed-Forward Models Can Outperform Recurrent Models</h1>
<p>Although it appears trainability and parallelization for feed-forward models  comes at the price of reduced accuracy, there have been several recent examples  showing that feed-forward networks can actually achieve the same accuracies as  their recurrent counterparts on benchmark tasks.</p>
<ul>
<li>
<p><strong>Language Modeling.</strong>  In language modeling, the goal is to predict the next word in a document given  all of the previous words. Feed-forward models make predictions using only the  $k$ most recent words, whereas recurrent models can potentially use the entire  document.  The <a href="https://arxiv.org/abs/1612.08083" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gated-Convolutional Language  Model</a> is a feed-forward autoregressive models  that is competitive with <a href="https://arxiv.org/abs/1602.02410" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">large LSTM baseline  models</a>. Despite using a truncation length of  $k=25$, the model outperforms a large LSTM on the  <a href="https://einstein.ai/research/the-wikitext-long-term-dependency-language-modeling-dataset" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wikitext-103</a>  benchmark, which is designed to reward models that capture long-term  dependencies. On the <a href="http://www.statmt.org/lm-benchmark/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Billion Word  Benchmark</a>, the model is slightly worse  than the largest LSTM, but is faster to train and uses fewer resources.</p>
</li>
<li>
<p><strong>Machine Translation.</strong>  The goal in machine translation is to map sequences of English words to  sequences of, say, French words. Feed-forward models make translations using  only $k$ words of the sentence, whereas recurrent models can leverage the entire  sentence.  Within the deep learning world, variants of the LSTM-based <a href="https://arxiv.org/abs/1409.0473" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sequence  to Sequence with Attention</a> model, particularly  <a href="https://arxiv.org/abs/1609.08144" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Google Neural Machine Translation</a>, were  superseded first by a fully <a href="https://arxiv.org/abs/1705.03122" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">convolutional sequence to  sequence</a> model and then by the  <a href="https://arxiv.org/abs/1706.03762" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Transformer</a>.<sup id="fnref:1"><a href="http://bair.berkeley.edu/blog/2018/08/06/recurrent/#fn:1" class="footnote" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">1</a></sup></p>
</li>
</ul>
<figure>
<p style="text-align:center;">          <img decoding="async" src="https://raw.githubusercontent.com/facebookresearch/fairseq/master/fairseq.gif" />      </p><figcaption>      <small>          Source: <a href="https://github.com/facebookresearch/fairseq/blob/master/fairseq.gif" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">           https://github.com/facebookresearch/fairseq/blob/master/fairseq.gif </a>      </small>      </figcaption></figure>
<ul>
<li>
<p><strong>Speech Synthesis.</strong>  In speech synthesis, one seeks to generate a realistic human speech signal.  Feed-forward models are limited to the past $k$ samples, whereas recurrent  models can use the entire history. Upon publication, the feed-forward,  autoregressive <a href="https://arxiv.org/abs/1609.03499" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">WaveNet</a> was a substantial  improvement over LSTM-RNN parametric models.</p>
</li>
<li>
<p><strong>Everthing Else.</strong>   Recently <a href="https://arxiv.org/abs/1803.01271" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Bai et al.</a> proposed a generic  feed-forward model leveraging dilated convolutions and showed it outperforms  recurrent baselines on tasks ranging from synthetic copying tasks to music  generation.</p>
</li>
</ul>
<h1 id="how-can-feed-forward-models-outperform-recurrent-ones">How Can Feed-Forward Models Outperform Recurrent Ones?</h1>
<p>In the examples above, feed-forward networks achieve results on par with or  better than recurrent networks. This is perplexing since recurrent models  seem to be more powerful a priori. One explanation for this phenomenon is  given by <a href="https://arxiv.org/abs/1612.08083" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Dauphin et al.</a>:</p>
<blockquote>
<p>The unlimited context offered by recurrent models is not strictly necessary  for language modeling.</p>
</blockquote>
<p>In other words, it’s possible you don’t need a large amount of context to do  well on the prediction task on average. <a href="https://arxiv.org/abs/1612.02526" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Recent theoretical  work</a> offers some evidence in favor of this view.</p>
<p>Another explanation is given by <a href="https://arxiv.org/abs/1803.01271" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Bai et al.</a>:</p>
<blockquote>
<p>The “infinite memory” advantage of RNNs is largely absent in practice.</p>
</blockquote>
<p>As Bai et al. report, even in experiments explicitly requiring long-term  context, RNN variants were unable to learn long sequences. On the Billion Word  Benchmark, an <a href="https://arxiv.org/abs/1703.10724" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">intriguing Google Technical  Report</a> suggests an LSTM $n$-gram model with  $n=13$ words of memory is as good as an LSTM with arbitrary context.</p>
<p>This evidence leads us to conjecture: <strong>Recurrent models <em>trained in practice</em>  are effectively feed-forward.</strong> This could happen either because truncated  backpropagation through time cannot learn patterns significantly longer than $k$  steps, or, more provocatively, because models <em>trainable by gradient descent</em>  cannot have long-term memory.</p>
<p>In <a href="https://arxiv.org/abs/1805.10369" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our recent paper</a>, we study the gap  between recurrent and feed-forward models trained using gradient descent. We  show if the recurrent model is <em>stable</em> (meaning the gradients can not explode),  then the model can be well-approximated by a feed-forward network for the  purposes of both <em>inference and training.</em> In other words, we show feed-forward  and stable recurrent models trained by gradient descent are <em>equivalent</em> in the  sense of making identical predictions at test-time.</p>
<p>Stability is a natural criterion for learnability of recurrent models. Outside  of the stable regime, gradient descent cannot be expected to work. Indeed, even  for very simple unstable models, gradient descent fails to converge to a  stationary point. While models trained in practice are not necessarily stable,  the performance of unstable models is likely in spite of, not due to, their  instability.</p>
<p>Using the <a href="https://einstein.ai/research/the-wikitext-long-term-dependency-language-modeling-dataset" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Wikitext-2 language modeling  benchmark</a>,  we show stability can be imposed on benchmark models without a loss in  performance. Concretely, we conducted a hyperparameter search to find the  best-performing unstable RNN and LSTM. Then, we re-trained both models while  enforcing the stability conditions derived in our paper. <strong>In both cases, the  unstable and stable models have the same test performance!</strong></p>
<table>
<thead>
<tr>
<th style="text-align: left">Recurrent Model</th>
<th style="text-align: center">Unstable (perplexity)</th>
<th style="text-align: center">Stable  (perplexity)</th>
</tr>
</thead>
<tbody>
<tr>
<td style="text-align: left">Tanh-RNN</td>
<td style="text-align: center">146.7</td>
<td style="text-align: center">143.5</td>
</tr>
<tr>
<td style="text-align: left">LSTM</td>
<td style="text-align: center">92.3</td>
<td style="text-align: center">95.1</td>
</tr>
</tbody>
</table>
<h1 id="conclusion">Conclusion</h1>
<p>Despite some initial attempts, there is still much to do to understand  why feed-forward models are competitive with recurrent ones and  shed light onto the trade-offs between sequence models. How much memory is  really needed to perform well on common sequence benchmarks? What are the  expressivity trade-offs between truncated RNNs (which can be considered  feed-forward) and the widely-used convolutional models?</p>
<p>Answering these questions is a step towards building a theory that can both  explain the strengths and limitations of our current methods and give guidance  about how to choose between different classes of models in concrete settings.</p>
<div class="footnotes">
<ol>
<li id="fn:1">
<p>The Transformer isn’t strictly a feed-forward model in the style described above (since it doesn’t make the $k$ step conditional independence assumption), but is not really a recurrent model because it doesn’t maintain a hidden state.&nbsp;</p>
<p>This article was initially published on the BAIR blog, and appears here with the authors’ permission.
</li>
</ol></div>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>One-shot imitation from watching videos</title>
		<link>https://robohub.org/one-shot-imitation-from-watching-videos/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Fri, 29 Jun 2018 22:12:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://robohub.org/one-shot-imitation-from-watching-videos/</guid>

					<description><![CDATA[<p>Learning a new skill by observing another individual, the ability to imitate, is
a key part of intelligence in human and animals. Can we enable a robot to do the
same, learning to manipulate a new object by simply watching a human
manipulating the object just as in the video below?</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/daml/demo_placing_peach.gif" height="360"><img src="http://bair.berkeley.edu/static/blog/daml/daml_placing_peach.gif" height="360"><br><i>
The robot learns to place the peach into the red bowl after watching the human
do so.
</i>
</p>

<!--more-->

<p>Such a capability would make it dramatically easier for us to communicate new
goals to robots &#8211; we could simply <em>show</em> robots what we want them to do, rather
than teleoperating the robot or engineering a reward function (an approach that
is difficult as it requires a full-fledged perception system). Many prior works
have investigated how well a robot can learn from an expert of its own kind
(i.e. through <a href="https://arxiv.org/abs/1710.04615" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">teleoperation</a> or <a href="https://ieeexplore.ieee.org/document/6249584/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">kinesthetic teaching</a>), which is usually
called <em><a href="http://bair.berkeley.edu/blog/2017/10/26/dart/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">imitation learning</a></em>. However, imitation learning of vision-based
skills usually requires a huge number of demonstrations of an expert performing
a skill. For example, a task like reaching toward a single fixed object using
raw pixel input requires 200 demonstrations to achieve good performance
according to <a href="https://arxiv.org/abs/1710.04615" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this prior work</a>. Hence a robot will struggle if there&#8217;s only
one demonstration presented.</p>

<p>Moreover, the problem becomes even more challenging when the robot needs to
imitate a human showing a certain manipulation skill. First, the robot arm looks
significantly different from the human arm. Second, engineering the right
correspondence between human demonstrations and robot demonstrations is
unfortunately extremely difficult. It&#8217;s not enough simple to track and remap the
motion: the task depends much more critically on how this motion affects objects
in the world, and we need a correspondence that is centrally based on the
interaction.</p>

<p>To enable the robot to imitate skills from one video of a human, we can allow it
to incorporate prior experience, rather than learn each skill completely from
scratch. By incorporating prior experience, the robot should also be able to
quickly learn to manipulate new objects while being invariant to shifts in
domain, such as a person providing a demonstration, a varying background scene,
or different viewpoint. We aim to achieve both of these abilities, few-shot
imitation and domain invariance, by learning to learn from demonstration data.
The technique, also called meta-learning and discussed in <a href="http://bair.berkeley.edu/blog/2017/07/18/learning-to-learn/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this previous blog
post</a>, is the key to how we equip robots with the ability to imitate by
observing a human.</p>

<h1>One-Shot Imitation Learning</h1>

<p>So how can we use meta-learning to make a robot quickly adapt to many different
objects? Our approach is to combine meta-learning with imitation learning to
enable one-shot imitation learning. The core idea is that provided a single
demonstration of a particular task, i.e. maneuvering a certain object, the robot
can quickly identify what the task is and successfully solve it under different
circumstances. <a href="https://arxiv.org/abs/1703.07326" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">A prior work</a> on one-shot imitation learning achieves
impressive results on simulated tasks such as block-stacking by learning to
learn across tens of thousands of demonstrations. If we want a physical robot to
able to emulate humans and manipulate a variety of novel objects, we need to
develop a new system that can learn to learn from demonstrations in the form of
videos using a dataset that can be practically collected in the real world.
First, we&#8217;ll discuss our approach for visual imitation of a single demonstration
collected via teleoperation. Then, we&#8217;ll show how it can be extended for
learning from videos of humans.</p>

<h2>One-Shot Visual Imitation Learning</h2>

<p>In order to make robots able to learn from watching videos, we combine imitation
learning with an efficient meta-learning algorithm, <a href="https://arxiv.org/abs/1703.03400" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">model-agnostic
meta-learning</a> (MAML). <a href="http://bair.berkeley.edu/blog/2017/07/18/learning-to-learn/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">This previous blog post</a> gives a nice overview of
the MAML algorithm. In this approach, we use a standard convolutional neural
network with parameters $\theta$ as our policy representation, mapping from an
image $o_t$ from the robot&#8217;s camera and the robot configuration $x_t$ (e.g.
joint angles and joint velocities) to robot actions $a_t$ (e.g. the linear and
angular velocity of the gripper) at time step $t$.</p>

<p>There are three main steps in this algorithm.</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/daml/mil_3_steps_diagram.png" width="600" alt="daml02"><br><i>
Three steps for our meta-learning algorithm.
</i>
</p>

<p>First, we collected a large dataset containing demonstrations of a teleoperated robot
performing many different tasks, which in our case, corresponds to manipulating
different objects. During the second step, we use MAML to learn an initial set
of policy parameters $\theta$, such that, after being provided a demonstration
for a certain object, we can run gradient descent with respect to the
demonstration to find a generalizable policy with parameters $\theta&#8217;$ for that
object. When using teleoperated demonstrations, the policy updates can be
computed by comparing the policy&#8217;s predicted action $\pi_\theta(o_t)$ to the
expert action :</p>

<p>Then, we optimize for the initial parameters $\theta$ by driving the updated
policy  to match the actions from another demonstration with
the same object. After meta-training, we can ask the robot to manipulate
completely unseen objects by computing gradient steps using a single
demonstration of that task. This step is called meta-testing.</p>

<p>As the method does not introduce any additional parameters for meta-learning and
optimization, it turns out to be quite data-efficient. Hence it can perform
various control tasks such as pushing and placing by just watching a
teleoperated robot demonstration:</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/daml/demo_robot_place.gif" height="360"><img src="http://bair.berkeley.edu/static/blog/daml/mil_robot_place.gif" height="360"><br><i>
Placing items into novel containers using a single demonstration. Left: demo.
Right: learned policy.
</i>
</p>

<h2>One-Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning</h2>

<p>The above method still relies on demonstrations coming from a teleoperated robot
rather than a human. To this end, we designed a domain-adaptive one-shot
imitation approach building on the above algorithm. We collected demonstrations
of many different tasks performed by both teleoperated robots <em>and</em> humans. Then, we
provide the human demonstration for computing the policy update and evaluate the
updated policy using a robot demonstration performing the same task. A diagram
illustrating this algorithm is below:</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/daml/daml_diagram.png" width="600"><br><i>
Overview of domain-adaptive meta-learning.
</i>
</p>

<p>Unfortunately, as a human demonstration is just a video of a human performing
the task, which doesn&#8217;t contain the expert actions , we can&#8217;t calculate
the policy update defined above. Instead, we propose to <em>learn</em> a loss function
for updating the policy, a loss function that doesn&#8217;t require action labels. The
intuition behind learning a loss function is that we can acquire a function that
only uses the available inputs, the unlabeled video, while still producing
gradients that are suitable for updating the policy parameters in a way that
produces a successful policy. While this might seem like an impossible task, it
is important to remember that the meta-training process still supervises the
policy with true robot actions after the gradient step.  The role of the learned
loss therefore may be interpreted as simply directing the parameter update to
modify the policy to pick up on the right visual cues in the scene, so that the
meta-trained action output will produce the right actions. We represent the
learned loss function using temporal convolutions, which can extract temporal
information in the video demonstration:</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/daml/temporal_conv.png" width="600"><br></p>

<p>We refer to this method as domain-adaptive meta-learning algorithm, as it learns
from data (e.g. videos of humans) from a different domain as the domain that the
robot&#8217;s policy operates in. Our method enables a PR2 robot to effectively learn
to push many different objects that are unseen during meta-training toward
target positions:</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/daml/push_obj3_demo.gif" height="280"><img src="http://bair.berkeley.edu/static/blog/daml/push_obj3_ours.gif" height="280"><br><i>
Learning to push a novel object by watching a human.
</i>
</p>

<p>and pick up many objects and place them onto target containers by watching a
human manipulates each object:</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/daml/pp2_demo.gif" height="360"><img src="http://bair.berkeley.edu/static/blog/daml/pp2_ours.gif" height="360"><br><i>
Learning to pick up a novel object and place it into a previously unseen bowl.
</i>
</p>

<p>We also evaluated the method using human demonstrations collected in a different
room with a different camera. The robot still performs these tasks reasonably
well:</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/daml/div_obj1_demo.gif" height="250"><img src="http://bair.berkeley.edu/static/blog/daml/div_obj1_bg0.gif" height="250"><br><i>
Learning to push a novel object by watching a human in a different environment
from a different viewpoint.
</i>
</p>

<h1>What&#8217;s Next?</h1>

<p>Now that we&#8217;ve taught a robot to learn to manipulate new objects by watching a
single video (which we also <a href="http://rail.eecs.berkeley.edu/nips_demo.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">demonstrated at NIPS 2017</a>), a natural next step
is to further scale these approaches to the setting where different tasks
correspond to entirely distinct motions and objectives, such as using a wide
variety of tools or playing a wide variety of sports. By considering
significantly more diversity in the underlying distribution of tasks, we hope
that these models will be able to achieve broader generalization, allowing
robots to quickly develop strategies for new situations. Further, the techniques
we developed here are not specific to robotic manipulation or even control. For
instance, both imitation learning and meta-learning have been used in the
context of language (examples <a href="http://papers.nips.cc/paper/5956-scheduled-sampling-for-sequence-prediction-with-recurrent-neural-networks" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a> and <a href="https://arxiv.org/abs/1803.02400" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a> respectively). In language
and other sequential decision-making settings, learning to imitate from a few
demonstrations is an interesting direction for future work.</p>

<hr><p>We would like to thank Sergey Levine and Pieter Abbeel for valuable feedback
when preparing this blog post.</p>

<p>This post is based on the following papers:</p>

<p><strong>One-Shot Visual Imitation Learning via Meta-Learning</strong><br>
Finn C., Yu T., Zhang T., Abbeel P., Levine S. CoRL 2017<br><a href="https://arxiv.org/abs/1709.04905" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a>, <a href="https://github.com/tianheyu927/mil" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">code</a>, <a href="https://sites.google.com/view/one-shot-imitation" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">videos</a></p>

<p><strong>One-Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning</strong><br>
Yu T., Finn C., Xie A., Dasari S., Zhang T., Abbeel P., Levine S. RSS 2018<br><a href="https://arxiv.org/abs/1802.01557" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a>, <a href="https://sites.google.com/view/daml" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">video</a></p>]]></description>
										<content:encoded><![CDATA[<p><strong>By Tianhe Yu and Chelsea Finn</strong></p>
<p>Learning a new skill by observing another individual, the ability to imitate, is a key part of intelligence in human and animals. Can we enable a robot to do the same, learning to manipulate a new object by simply watching a human manipulating the object just as in the video below?</p>
<p> <span id="more-104329"></span></p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/demo_placing_peach.gif" height="360" style="margin: 10px;" /> <img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/daml_placing_peach.gif" height="360" style="margin: 10px;" /> <br /> <i> The robot learns to place the peach into the red bowl after watching the human do so. </i> </p>
<p>  <!--more-->  </p>
<p>Such a capability would make it dramatically easier for us to communicate new goals to robots – we could simply <em>show</em> robots what we want them to do, rather than teleoperating the robot or engineering a reward function (an approach that is difficult as it requires a full-fledged perception system). Many prior works have investigated how well a robot can learn from an expert of its own kind (i.e. through <a href="https://arxiv.org/abs/1710.04615" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">teleoperation</a> or <a href="https://ieeexplore.ieee.org/document/6249584/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">kinesthetic teaching</a>), which is usually called <em><a href="http://bair.berkeley.edu/blog/2017/10/26/dart/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">imitation learning</a></em>. However, imitation learning of vision-based skills usually requires a huge number of demonstrations of an expert performing a skill. For example, a task like reaching toward a single fixed object using raw pixel input requires 200 demonstrations to achieve good performance according to <a href="https://arxiv.org/abs/1710.04615" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this prior work</a>. Hence a robot will struggle if there’s only one demonstration presented.</p>
<p>Moreover, the problem becomes even more challenging when the robot needs to imitate a human showing a certain manipulation skill. First, the robot arm looks significantly different from the human arm. Second, engineering the right correspondence between human demonstrations and robot demonstrations is unfortunately extremely difficult. It’s not enough simple to track and remap the motion: the task depends much more critically on how this motion affects objects in the world, and we need a correspondence that is centrally based on the interaction.</p>
<p>To enable the robot to imitate skills from one video of a human, we can allow it to incorporate prior experience, rather than learn each skill completely from scratch. By incorporating prior experience, the robot should also be able to quickly learn to manipulate new objects while being invariant to shifts in domain, such as a person providing a demonstration, a varying background scene, or different viewpoint. We aim to achieve both of these abilities, few-shot imitation and domain invariance, by learning to learn from demonstration data. The technique, also called meta-learning and discussed in <a href="http://bair.berkeley.edu/blog/2017/07/18/learning-to-learn/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this previous blog post</a>, is the key to how we equip robots with the ability to imitate by observing a human.</p>
<h1 id="one-shot-imitation-learning">One-Shot Imitation Learning</h1>
<p>So how can we use meta-learning to make a robot quickly adapt to many different objects? Our approach is to combine meta-learning with imitation learning to enable one-shot imitation learning. The core idea is that provided a single demonstration of a particular task, i.e. maneuvering a certain object, the robot can quickly identify what the task is and successfully solve it under different circumstances. <a href="https://arxiv.org/abs/1703.07326" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">A prior work</a> on one-shot imitation learning achieves impressive results on simulated tasks such as block-stacking by learning to learn across tens of thousands of demonstrations. If we want a physical robot to able to emulate humans and manipulate a variety of novel objects, we need to develop a new system that can learn to learn from demonstrations in the form of videos using a dataset that can be practically collected in the real world. First, we’ll discuss our approach for visual imitation of a single demonstration collected via teleoperation. Then, we’ll show how it can be extended for learning from videos of humans.</p>
<h2 id="one-shot-visual-imitation-learning">One-Shot Visual Imitation Learning</h2>
<p>In order to make robots able to learn from watching videos, we combine imitation learning with an efficient meta-learning algorithm, <a href="https://arxiv.org/abs/1703.03400" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">model-agnostic meta-learning</a> (MAML). <a href="http://bair.berkeley.edu/blog/2017/07/18/learning-to-learn/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">This previous blog post</a> gives a nice overview of the MAML algorithm. In this approach, we use a standard convolutional neural network with parameters $\theta$ as our policy representation, mapping from an image $o_t$ from the robot’s camera and the robot configuration $x_t$ (e.g. joint angles and joint velocities) to robot actions $a_t$ (e.g. the linear and angular velocity of the gripper) at time step $t$.</p>
<p>There are three main steps in this algorithm.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/mil_3_steps_diagram.png" width="600" alt="daml02" /><br /> <i> Three steps for our meta-learning algorithm. </i> </p>
<p>First, we collected a large dataset containing demonstrations of a teleoperated robot performing many different tasks, which in our case, corresponds to manipulating different objects. During the second step, we use MAML to learn an initial set of policy parameters $\theta$, such that, after being provided a demonstration for a certain object, we can run gradient descent with respect to the demonstration to find a generalizable policy with parameters $\theta’$ for that object. When using teleoperated demonstrations, the policy updates can be computed by comparing the policy’s predicted action $\pi_\theta(o_t)$ to the expert action <script type="math/tex">a^*_t</script>:</p>
<p>  <script type="math/tex; mode=display">\theta’ \leftarrow \theta - \alpha \nabla_\theta \sum_t || \pi_\theta(o_t) - a^*_t || ^2.</script>  </p>
<p>Then, we optimize for the initial parameters $\theta$ by driving the updated policy <script type="math/tex">\pi_{\theta’}</script> to match the actions from another demonstration with the same object. After meta-training, we can ask the robot to manipulate completely unseen objects by computing gradient steps using a single demonstration of that task. This step is called meta-testing.</p>
<p>As the method does not introduce any additional parameters for meta-learning and optimization, it turns out to be quite data-efficient. Hence it can perform various control tasks such as pushing and placing by just watching a teleoperated robot demonstration:</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/demo_robot_place.gif" height="360" style="margin: 10px;" /> <img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/mil_robot_place.gif" height="360" style="margin: 10px;" /> <br /> <i> Placing items into novel containers using a single demonstration. Left: demo. Right: learned policy. </i> </p>
<h2 id="one-shot-imitation-from-observing-humans-via-domain-adaptive-meta-learning">One-Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning</h2>
<p>The above method still relies on demonstrations coming from a teleoperated robot rather than a human. To this end, we designed a domain-adaptive one-shot imitation approach building on the above algorithm. We collected demonstrations of many different tasks performed by both teleoperated robots <em>and</em> humans. Then, we provide the human demonstration for computing the policy update and evaluate the updated policy using a robot demonstration performing the same task. A diagram illustrating this algorithm is below:</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/daml_diagram.png" width="600" /><br /> <i> Overview of domain-adaptive meta-learning. </i> </p>
<p>Unfortunately, as a human demonstration is just a video of a human performing the task, which doesn’t contain the expert actions <script type="math/tex">a^*_t</script>, we can’t calculate the policy update defined above. Instead, we propose to <em>learn</em> a loss function for updating the policy, a loss function that doesn’t require action labels. The intuition behind learning a loss function is that we can acquire a function that only uses the available inputs, the unlabeled video, while still producing gradients that are suitable for updating the policy parameters in a way that produces a successful policy. While this might seem like an impossible task, it is important to remember that the meta-training process still supervises the policy with true robot actions after the gradient step.  The role of the learned loss therefore may be interpreted as simply directing the parameter update to modify the policy to pick up on the right visual cues in the scene, so that the meta-trained action output will produce the right actions. We represent the learned loss function using temporal convolutions, which can extract temporal information in the video demonstration:</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/temporal_conv.png" width="600" /> </p>
<p>We refer to this method as domain-adaptive meta-learning algorithm, as it learns from data (e.g. videos of humans) from a different domain as the domain that the robot’s policy operates in. Our method enables a PR2 robot to effectively learn to push many different objects that are unseen during meta-training toward target positions:</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/push_obj3_demo.gif" height="280" style="margin: 10px;" /> <img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/push_obj3_ours.gif" height="280" style="margin: 10px;" /> <br /> <i> Learning to push a novel object by watching a human. </i> </p>
<p>and pick up many objects and place them onto target containers by watching a human manipulates each object:</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/pp2_demo.gif" height="360" style="margin: 10px;" /> <img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/pp2_ours.gif" height="360" style="margin: 10px;" /> <br /> <i> Learning to pick up a novel object and place it into a previously unseen bowl. </i> </p>
<p>We also evaluated the method using human demonstrations collected in a different room with a different camera. The robot still performs these tasks reasonably well:</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/div_obj1_demo.gif" height="500" width="500"  style="margin: 10px;" />  </p>
<p><img decoding="async" src="http://bair.berkeley.edu/static/blog/daml/div_obj1_bg0.gif" height="250" style="margin: 10px;" /> <br /> <i> Learning to push a novel object by watching a human in a different environment from a different viewpoint. </i> </p>
<h1 id="whats-next">What’s Next?</h1>
<p>Now that we’ve taught a robot to learn to manipulate new objects by watching a single video (which we also <a href="http://rail.eecs.berkeley.edu/nips_demo.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">demonstrated at NIPS 2017</a>), a natural next step is to further scale these approaches to the setting where different tasks correspond to entirely distinct motions and objectives, such as using a wide variety of tools or playing a wide variety of sports. By considering significantly more diversity in the underlying distribution of tasks, we hope that these models will be able to achieve broader generalization, allowing robots to quickly develop strategies for new situations. Further, the techniques we developed here are not specific to robotic manipulation or even control. For instance, both imitation learning and meta-learning have been used in the context of language (examples <a href="http://papers.nips.cc/paper/5956-scheduled-sampling-for-sequence-prediction-with-recurrent-neural-networks" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a> and <a href="https://arxiv.org/abs/1803.02400" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a> respectively). In language and other sequential decision-making settings, learning to imitate from a few demonstrations is an interesting direction for future work.</p>
<hr />
<p>We would like to thank Sergey Levine and Pieter Abbeel for valuable feedback when preparing this blog post. This article was initially published on the <a href="http://bair.berkeley.edu/blog/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">BAIR blog</a>, and appears here with the authors’ permission. </p>
<p>This post is based on the following papers:</p>
<p><strong>One-Shot Visual Imitation Learning via Meta-Learning</strong><br /> Finn C.<script type="math/tex">^*</script>, Yu T.<script type="math/tex">^*</script>, Zhang T., Abbeel P., Levine S. CoRL 2017<br /> <a href="https://arxiv.org/abs/1709.04905" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a>, <a href="https://github.com/tianheyu927/mil" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">code</a>, <a href="https://sites.google.com/view/one-shot-imitation" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">videos</a></p>
<p><strong>One-Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning</strong><br /> Yu T.<script type="math/tex">^*</script>, Finn C.<script type="math/tex">^*</script>, Xie A., Dasari S., Zhang T., Abbeel P., Levine S. RSS 2018<br /> <a href="https://arxiv.org/abs/1802.01557" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">paper</a>, <a href="https://sites.google.com/view/daml" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">video</a></p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>BDD100K: A large-scale diverse driving video database</title>
		<link>https://robohub.org/bdd100k-a-large-scale-diverse-driving-video-database/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Fri, 01 Jun 2018 01:42:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://robohub.org/bdd100k-a-large-scale-diverse-driving-video-database/</guid>

					<description><![CDATA[TL;DR, we released the largest and most diverse driving video dataset with rich
annotations called BDD100K. You can access the data for research now at http://bdd-data.berkeley.edu.  We  have
recently released an arXiv
report on it. And there is still ...]]></description>
										<content:encoded><![CDATA[<p><strong>By Fisher Yu</strong></p>
<p>TL;DR, we released the largest and most diverse driving video dataset with richannotations called BDD100K. You can access the data for research now at <a href="http://bdd-data.berkeley.edu/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">http://bdd-data.berkeley.edu</a>.  We  haverecently released <a href="https://arxiv.org/abs/1805.04687" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">an arXivreport</a> on it. And there is still time to participate in <a href="http://bdd-data.berkeley.edu/wad-2018.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our CVPR 2018 challenges</a>!</p>
<p><span id="more-102710"></span></p>
<div class="keep-aspect"><iframe title="BDD100K Trailer" width="500" height="281" src="https://www.youtube-nocookie.com/embed/IGi9K9FY35Y?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></div>
<p></p>
<h2 id="large-scale-diverse-driving-video-pick-four">Large-scale, Diverse, Driving, Video: Pick Four</h2>
<p>Autonomous driving is poised to change the life in every community. However,recent events show that it is not clear yet how a man-made perception system canavoid even seemingly obvious mistakes when a driving system is deployed in thereal world. As computer vision researchers, we are interested in exploring thefrontiers of perception algorithms for self-driving to make it safer. To designand test potential algorithms, we would like to make use of all the informationfrom the data collected by a real driving platform. Such data has four majorproperties: it is large-scale, diverse, captured on the street, and withtemporal information. Data diversity is especially important to test therobustness of perception algorithms. However, current open datasets can onlycover a subset of the properties described above. Therefore, with the help of <a href="https://www.getnexar.com/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Nexar</a>, we are releasing the BDD100Kdatabase, which is the largest and most diverse open driving video dataset sofar for computer vision research. This project is organized and sponsored by <a href="https://deepdrive.berkeley.edu/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Berkeley DeepDrive</a> IndustryConsortium, which investigates state-of-the-art technologies in computer visionand machine learning for automotive applications.</p>
<p style="text-align:center;"><img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/bdd/geo_distribution.jpg" /><br /><i>Locations of a random video subset.</i></p>
<p>As suggested in the name, our dataset consists of 100,000 videos. Each video isabout 40 seconds long, 720p, and 30 fps. The videos also come with GPS/IMUinformation recorded by cell-phones to show rough driving trajectories. Ourvideos were collected from diverse locations in the United States, as shown inthe figure above. Our database covers different weather conditions, includingsunny, overcast, and rainy, as well as  different times of day including daytimeand nighttime. The table below summarizes comparisons with previous datasets,which shows our dataset is much larger and more diverse.</p>
<p><!--

<table>     

<tr>         

<th></th>

        

 

<th style="text-align: center"> <a href="http://www.cvlibs.net/publications/Geiger2012CVPR.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"> KITTI</th>

         

<th style="text-align: center"> <a href="https://arxiv.org/abs/1604.01685"> Cityscapes</th>

         

<th style="text-align: center"> <a href="https://arxiv.org/pdf/1803.06184v1.pdf"> ApolloScape</th>

         

<th style="text-align: center"> <a href="https://research.mapillary.com/img/publications/ICCV17a.pdf"> Mapillary</th>

         

<th style="text-align: center"> <a href="https://arxiv.org/abs/1805.04687"> BDD100K </a> </th>

     </tr>

     

<tr>         

<td align="center"># Sequences</td>

         

<td align="center">22</td>

         

<td align="center">~50</td>

         

<td align="center">4</td>

         

<td align="center">N/A</td>

         

<td align="center">100,000</td>

    </tr>

    

<tr>         

<td align="center"># Images</td>

         

<td align="center">14,999</td>

         

<td align="center">5000 (+2000)</td>

         

<td align="center">143,906</td>

         

<td align="center">25,000</td>

         

<td align="center">120,000,000</td>

    </tr>

    

<tr>         

<td align="center">Multiple Cities</td>

         

<td align="center" style="color:red">No</td>

         

<td align="center" style="color:green">Yes</td>

         

<td align="center" style="color:red">No</td>

         

<td align="center" style="color:green">Yes</td>

         

<td align="center" style="color:green">Yes</td>

    </tr>

         

<tr>         

<td align="center">Multiple Weathers</td>

         

<td align="center" style="color:red">No</td>

         

<td align="center" style="color:red">No</td>

         

<td align="center" style="color:red">No</td>

         

<td align="center" style="color:green">Yes</td>

         

<td align="center" style="color:green">Yes</td>

    </tr>

         

<tr>         

<td align="center">Multiple Times of Day</td>

         

<td align="center" style="color:red">No</td>

         

<td align="center" style="color:red">No</td>

         

<td align="center" style="color:red">No</td>

         

<td align="center" style="color:green">Yes</td>

         

<td align="center" style="color:green">Yes</td>

    </tr>

         

<tr>         

<td align="center">Multiple Scene types</td>

         

<td align="center" style="color:green">Yes</td>

         

<td align="center" style="color:red">No</td>

         

<td align="center" style="color:red">No</td>

         

<td align="center" style="color:green">Yes</td>

         

<td align="center" style="color:green">Yes</td>

    </tr>

</table>

--></p>
<p style="text-align:center;"><img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/bdd/table0.png" /><br /><i>Comparisons with some other street scene datasets. It is hard to fairly compare# images between datasets, but we list them here as a rough reference.</i></p>
<p>The videos and their trajectories can be useful for imitation learning ofdriving policies, as in our <a href="https://arxiv.org/abs/1612.01079" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">CVPR 2017paper</a>. To facilitate computer vision research on our large-scale dataset, wealso provide basic annotations on the video keyframes, as detailed in the nextsection. You can download the data and annotations now at <a href="http://bdd-data.berkeley.edu/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">http://bdd-data.berkeley.edu</a>.</p>
<h2 id="annotations">Annotations</h2>
<p>We sample a keyframe at the 10th second from each video and provide annotationsfor those keyframes. They are labeled at several levels: image tagging, roadobject bounding boxes, drivable areas, lane markings, and full-frame instancesegmentation. These annotations will help us understand the diversity of thedata and object statistics in different types of scenes. We will discuss thelabeling process in a different blog post. More information about theannotations can be found in our <a href="https://arxiv.org/abs/1805.04687" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">arXivreport</a>.</p>
<p style="text-align:center;"><img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/bdd/annotation_examples.png" /><br /><i>Overview of our annotations.</i></p>
<h3 id="road-object-detection">Road Object Detection</h3>
<p>We label object bounding boxes for objects that commonly appear on the road onall of the 100,000 keyframes to understand the distribution of the objects andtheir locations. The bar chart below shows the object counts. There are alsoother ways to play with the statistics in our annotations. For example, we cancompare the object counts under different weather conditions or in differenttypes of scenes. This chart also shows the diverse set of objects that appear inour dataset, and the scale of our dataset –  more than 1 million cars. Thereader should be reminded here that those are distinct objects with distinctappearances and contexts.</p>
<p style="text-align:center;"><img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/bdd/bbox_instance.png" /><br /><i>Statistics of different types of objects.</i></p>
<p>Our dataset is also suitable for studying some particular domains. For example,if you are interested in detecting and avoiding pedestrians on the streets, youalso have a reason to study our dataset since it contains more pedestrianinstances than previous specialized datasets as shown in the table below.</p>
<p><!--

<table>     

<tr>         

<th></th>

         

<th style="text-align: center"> <a href="https://core.ac.uk/download/pdf/4875878.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"> Caltech</th>

         

<th style="text-align: center"> <a href='http://www.cvlibs.net/publications/Geiger2012CVPR.pdf'> KITTI</th>

         

<th style="text-align: center"> <a href='https://arxiv.org/abs/1702.05693'> CityPerson</th>

         

<th style="text-align: center"> <a href='https://arxiv.org/abs/1805.04687'> BDD100K </a> </th>

     </tr>

     

<tr>         

<td align="center"># persons</td>

         

<td align="center">1,273</td>

         

<td align="center">6,336</td>

         

<td align="center">19,654</td>

         

<td align="center">86,047</td>

    </tr>

    

<tr>         

<td align="center"># per image</td>

         

<td align="center">1.4</td>

         

<td align="center">0.8</td>

         

<td align="center">7.0</td>

         

<td align="center">1.2</td>

    </tr>

</table>

--></p>
<p style="text-align:center;"><img decoding="async" width="600" src="http://bair.berkeley.edu/static/blog/bdd/table1.png" /><br /><i>Comparisons with other pedestrian datasets regarding training set size.</i></p>
<h3 id="lane-markings">Lane Markings</h3>
<p>Lane markings are important road instructions for human drivers. They are alsocritical cues of driving direction and localization for the autonomous drivingsystems when GPS or maps does not have accurate global coverage. We divide thelane markings into two types based on how they instruct the vehicles in thelanes. Vertical lane markings (marked in red in the figures below) indicatemarkings that are  along the driving direction of their lanes. Parallel lanemarkings (marked in blue in the figures below) indicate those that are  for thevehicles in the lanes to stop. We also provide attributes for the markings suchas solid vs. dashed and double vs. single.</p>
<p style="text-align:center;"><img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/bdd/lane_markings.png" /></p>
<p>If you are ready to try out your lane marking prediction algorithms, please lookno further. Here is the comparison with existing lane marking datasets.</p>
<p><!--

<table>     

<tr>         

<th></th>

         

<th style="text-align: center"> Training </th>

         

<th style="text-align: center"> Total </th>

         

<th style="text-align: center"> Sequences </th>

         

<th style="text-align: center"> Weather </th>

         

<th style="text-align: center"> Time </a> </th>

         

<th style="text-align: center"> Attributes </a> </th>

     </tr>

     

<tr>         

<td><a href="https://arxiv.org/abs/1411.7113" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"> Caltech Lanes Dataset</a></td>

         

<td align="center">-</td>

         

<td align="center">1,224</td>

         

<td align="center">4</td>

         

<td align="center">1</td>

         

<td align="center">1</td>

         

<td align="center">2</td>

    </tr>

    

<tr>       

<td><a href="https://ieeexplore.ieee.org/document/6232144/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"> Road Marking Dataset</a></td>

       

<td align="center">-</td>

       

<td align="center">1,443</td>

       

<td align="center">29</td>

       

<td align="center">2</td>

       

<td align="center">3</td>

       

<td align="center">10</td>

    </tr>

    

<tr>       

<td><a href="http://www.cvlibs.net/projects/autonomous_vision_survey/literature/Fritsch2013ITSC.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"> KITTI Road</a></td>

       

<td align="center">289</td>

       

<td align="center">579</td>

       

<td align="center">-</td>

       

<td align="center">1</td>

       

<td align="center">1</td>

       

<td align="center">2</td>

   </tr>

   

<tr>       

<td><a href="https://arxiv.org/abs/1710.06288" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"> VPGNet</a></td>

       

<td align="center">14,783</td>

       

<td align="center">21,097</td>

       

<td align="center">-</td>

       

<td align="center">4</td>

       

<td align="center">2</td>

       

<td align="center">17</td>

   </tr>

   

<tr>       

<td><a href="https://arxiv.org/abs/1805.04687" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"> BDD100K</a></td>

       

<td align="center">70,000</td>

       

<td align="center">100,000</td>

       

<td align="center">100,000</td>

       

<td align="center">6</td>

       

<td align="center">3</td>

       

<td align="center">11</td>

   </tr>

</table>

--></p>
<p style="text-align:center;"><img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/bdd/table2.png" /></p>
<h3 id="drivable-areas">Drivable Areas</h3>
<p>Whether we can drive on a road does not only depend on lane markings and trafficdevices. It also depends on the complicated interactions with other objectssharing the road. In the end, it  is important to understand which area can bedriven on. To investigate this problem, we also provide segmentation annotationsof drivable areas as shown below. We divide  the drivable areas into twocategories based on the trajectories of the ego vehicle: direct drivable, andalternative drivable. Direct drivable, marked in  red, means the ego vehicle hasthe road priority and can keep driving in that area. Alternative drivable,marked in  blue, means the ego vehicle can drive in the area, but has to becautious since the road priority  potentially belongs to other vehicles.</p>
<p style="text-align:center;"><img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/bdd/drivable_area.png" /></p>
<h3 id="full-frame-segmentation">Full-frame Segmentation</h3>
<p>It has been shown on Cityscapes dataset that full-frame fine instancesegmentation can greatly bolster research in dense prediction and objectdetection, which are pillars of a wide range of computer vision applications. Asour videos are in a different domain, we provide instance segmentationannotations as well to compare the domain shift relative by different datasets.It can be expensive and laborious to obtain full pixel-level segmentation.Fortunately, with our own labeling tool, the labeling cost could be reduced by50%. In the end, we label a subset of 10K images with full-frame instancesegmentation. Our label set is compatible with the training annotations inCityscapes to make it easier to study domain shift between the datasets.</p>
<p style="text-align:center;"><img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/bdd/segmentation.jpg" /></p>
<h2 id="driving-challenges">Driving Challenges</h2>
<p>We are hosting <a href="http://bdd-data.berkeley.edu/wad-2018.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">threechallenges</a> in CVPR 2018 Workshop on Autonomous Driving based on our data:road object detection, drivable area prediction, and domain adaptation ofsemantic segmentation. The detection task requires your algorithm to find all ofthe target objects in our testing images and drivable area prediction requiressegmenting the areas a car can drive in. In domain adaptation, the testing datais collected in China. Systems are thus challenged to get models learned in theUS to work in the crowded streets in Beijing, China. You can submit your resultsnow after <a href="http://bdd-data.berkeley.edu/login.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">logging in ouronline submission portal</a>. Make sure to check out <a href="https://github.com/ucbdrive/bdd-data" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our toolkit</a> to jump start yourparticipation.</p>
<p>Join our CVPR workshop challenges to claim your cash prizes!!!</p>
<h2 id="future-work">Future Work</h2>
<p>The perception system for self-driving is by no means only about monocularvideos. It may also include panorama and stereo videos as well as  other typesof sensors like LiDAR and radar. We hope to provide and study thosemulti-modality sensor data as well in the near future.</p>
<h2 id="reference-links">Reference Links</h2>
<p><a href="https://core.ac.uk/download/pdf/4875878.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"> Caltech,<a href="http://www.cvlibs.net/publications/Geiger2012CVPR.pdf"> KITTI,<a href="https://arxiv.org/abs/1702.05693"> CityPerson,<a href="https://arxiv.org/abs/1604.01685"> Cityscapes,<a href="https://arxiv.org/pdf/1803.06184v1.pdf"> ApolloScape,<a href="https://research.mapillary.com/img/publications/ICCV17a.pdf"> Mapillary,<a href="https://arxiv.org/abs/1411.7113"> Caltech Lanes Dataset,<a href="https://ieeexplore.ieee.org/document/6232144/"> Road Marking Dataset,<a href="http://www.cvlibs.net/projects/autonomous_vision_survey/literature/Fritsch2013ITSC.pdf"> KITTI Road,<a href="https://arxiv.org/abs/1710.06288"> VPGNet</a></a></a></a></a></a></a></a></a></a></p>
<p>This article was initially published on the <a href="“http://bair.berkeley.edu/blog/2017/06/20/welcome/”" data-wpel-link="internal">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>TDM: From model-free to model-based deep reinforcement learning</title>
		<link>https://robohub.org/tdm-from-model-free-to-model-based-deep-reinforcement-learning/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Fri, 01 Jun 2018 01:35:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://robohub.org/tdm-from-model-free-to-model-based-deep-reinforcement-learning/</guid>

					<description><![CDATA[<p>
You&#8217;ve decided that you want to bike from your house by UC Berkeley to the
Golden Gate Bridge. It&#8217;s a nice 20 mile ride, but there&#8217;s a problem: you&#8217;ve
never ridden a bike before! To make matters worse, you are new to the Bay Area,
and all you have is a good ol&#8217; fashion map to guide you. How do you get started?
</p>

<p>
Let&#8217;s first figure out how to ride a bike. One strategy would be to do a lot of
studying and planning. Read books on how to ride bicycles. Study physics and
anatomy. Plan out all the different muscle movements that you&#8217;ll make in
response to each perturbation. This approach is noble, but for anyone who&#8217;s ever
learned to ride a bike, they know that this strategy is doomed to fail. There&#8217;s
only one way to learn how to ride a bike: trial and error. Some tasks like
riding a bike are just too complicated to plan out in your head.  </p>

<p>
Once you&#8217;ve learned how to ride your bike, how would you get to the Golden Gate
Bridge? You could reuse your trial-and-error strategy. Take a few random turns
and see if you end up at the Golden Gate Bridge. Unfortunately, this strategy
would take a very, very long time. For this sort of problem, planning is a much
faster strategy, and requires considerably less real-world experience and
trial-and-error. In reinforcement learning terms, it is more
<i>sample-efficient</i>.
</p>

&#60;!--
<p style="text-align:center">
<img src="http://bair.berkeley.edu/static/blog/tdm/riding-bike-small.png" alt="Some skills you learn by trial and error." /><br />
<i>
Some skills you learn by trial and error.
</i>
</p>

<p style="text-align:center">
<img src="http://bair.berkeley.edu/static/blog/tdm/hitchhiker-small.png" alt="Other times, planning ahead is better." /><br />
<i>
Other times, planning ahead is better.
</i>
</p>
--&#62;

<p>
</p><table><tr><td>
      <img src="http://bair.berkeley.edu/static/blog/tdm/riding-bike-small.png" height="260"></td>
    <td>
      <img src="http://bair.berkeley.edu/static/blog/tdm/hitchhiker-small.png" height="260"></td>
  </tr></table><p>
<i>
Left: some skills you learn by trial and error. Right: other times, planning
ahead is better.
</i>
</p>

<p>
While simple, this thought experiment highlights some important aspects of human
intelligence. For some tasks, we use a trial-and-error approach, and for others
we use a planning approach. A similar phenomena seems to have emerged in
reinforcement learning (RL). In the parlance of RL, empirical results show that
some tasks are better suited for model-free (trial-and-error) approaches, and
others are better suited for model-based (planning) approaches.  </p>

<p>
However, the biking analogy also highlights that the two systems are not
completely independent. In particularly, to say that learning to ride a bike is
<i>just</i> trial-and-error is an oversimplification. In fact, when learning to
bike by trial-and-error, you&#8217;ll employ a bit of planning. Perhaps your plan will
initially be, &#8220;Don&#8217;t fall over.&#8221; As you improve, you&#8217;ll make more ambitious
plans, such as, &#8220;Bike forwards for two meters without falling over.&#8221; Eventually,
your bike-riding skills will be so proficient that you can start to plan in very
abstract terms (&#8220;Bike to the end of the road.&#8221;) to the point that all there is
left to do is planning and you no longer need to worry about the nitty-gritty
details of riding a bike. We see that there is a gradual transition from the
model-free (trial-and-error) strategy to a model-based (planning) strategy. If
we could develop artificial intelligence algorithms--and specifically RL
algorithms--that mimic this behavior, it could result in an algorithm that both
performs well (by using trial-and-error methods early on) and is sample
efficient (by later switching to a planning approach to achieve more abstract
goals).
</p>

<p>
This post covers temporal difference model (TDM), which is a RL algorithm that
captures this smooth transition between model-free and model-based RL. Before
describing TDMs, we start by first describing how a typical model-based RL
algorithm works.
</p>

<!--more-->

<h2>Model-Based Reinforcement Learning</h2>

<p>
In reinforcement learning, we have some some state space $\mathcal{S}$ and
action space $\mathcal{A}$. If at time $t$ we are in state $s_t \in \mathcal{S}$
and take action $a_t\in \mathcal{A}$, we transition to a new state $s_{t+1} =
f(s_t, a_t)$ according to a dynamics model $f: \mathcal{S} \times \mathcal{A}
\mapsto \mathcal{S}$. The goal is to maximize rewards summed over the visited
state: $\sum_{t=1}^{T-1} r(s_t, a_, s_{t+1})$. Model-based RL algorithms assume
you are given (or learn) the dynamics model $f$. Given this dynamics model,
there are a variety of model-based algorithms. For this post, we consider
methods that perform the following optimization to choose a sequence of actions
and states to maximize rewards: 
</p>

<p>
  The optimization says to choose a sequence of states and actions that you maximize the rewards, while ensuring that the trajectory is feasible. Here, feasible means that each state-action-next-state transition is valid. For example, in the image below if you start in state $s_t$ and take action $a_t$, only the top $s_{t+1}$ results in a feasible transition.
</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/tdm/bike-feasibility-compressed.png" alt="Planning a trip to the Golden Gate Bridge would be much easier if you could defy physics. However, the constraint in the model-based optimization problem ensures that only trajectories like the top row will be outputted. The bottom two trajectories may have high reward, but they&#8217;re not feasible." width="80%"><br><i>
Planning a trip to the Golden Gate Bridge would be much easier if you could defy
physics. However, the constraint in the model-based optimization problem ensures
that only trajectories like the top row will be outputted. The bottom two
trajectories may have high reward, but they&#8217;re not feasible.
</i>
</p>

<p>
In our biking problem, the optimization might result in a biking plan from
Berkeley (top right) to the Golden Gate Bridge (middle left) that looks like
this:
</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/tdm/mb-bike-plan-small.png" alt="An example of a plan (states and actions) outputted the optimization problem." width="80%"><br><i>
An example of a plan (states and actions) outputted the optimization problem.
</i>
</p>

<p>
While conceptually nice, this plan is not very realistic. Model-based approaches
use a model $f(s, a)$ that predict the state at the very next time step. In
robotics, a time step usually corresponds to a tenth or a hundredth of a second.
So perhaps a more realistic depiction of the resulting plan might look like:
</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/tdm/mb-short-time-small.png" alt="A more realistic plan." width="80%"><br><i>
A more realistic plan.
</i>
</p>

<p>
If we think about how we plan in everyday life, we realize that we plan at much
more temporally abstract terms. Rather than planning the position that our bike
will be at the next tenth of a second, we make longer-term plan things like, &#8220;I
will go to the end of the road.&#8221; Furthermore, we can only make these temporally
abstract plans once we&#8217;ve learned how to ride a bike in the first place. As
discussed earlier, we need some way to (1) start the learning using a
trial-and-error approach and (2) provide a mechanism to gradually increase the
level of abstraction that we use to plan. For this, we introduce temporal
difference models.
</p>

<h2>Temporal Difference Models</h2>
<p>
  A temporal difference model (TDM)$^\dagger$, which we will write as $Q(s, a, s_g, \tau)$, is a function that, given a state $s \in \mathcal{S}$, action $a \in \mathcal{A}$, and goal state $s_g \in \mathcal{S}$, predicts how close an agent can get to the goal within $\tau$ time steps. Intuitively, a TDM answers the question, &#8220;If I try to bike to San Francisco in 30 minutes, how close will I get?&#8221; For robotics, a natural way to measure closeness is use Euclidean distance.
</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/tdm/tdm-visualization-small.png" alt="A TDM predicts how close you will get to the goal (Golden Gate Bridge) after a fixed amount of time. After 30 minutes of biking, maybe you only reach the grey biker in the image above. In this case, the grey line represents the distance that the TDM should predict." width="80%"><br><i>
A TDM predicts how close you will get to the goal (Golden Gate Bridge) after a
fixed amount of time. After 30 minutes of biking, maybe you only reach the grey
biker in the image above. In this case, the grey line represents the distance
that the TDM should predict.
</i>
</p>

<p>
For those familiar with reinforcement learning, it turns out that a TDM can be
viewed as a goal-conditioned Q function in a finite-horizon MDP. Because a TDM
is just another Q function, we can train it with model-free (trial-and-error)
algorithms. We use <a href="https://arxiv.org/abs/1509.02971" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">deep deterministic
policy gradient</a> (DDPG) to train a TDM and retroactively relabel the goal and
time horizon to increase the sample efficiency of our learning algorithm. In
theory, any Q-learning algorithm could be used to train the TDM, but we found
this to be effective. We encourage readers to check out the paper for more
details.
</p>

<h3>Planning with a TDM</h3>
<p>
Once we train a TDM, how can we use it to plan? It turns out that we can plan with the following optimization:
</p>

<p>
The intuition is similar to the model-based formulation. Choose a sequence of
actions and states that maximize rewards and that are feasible. A key difference
is that we only plan <i>every $K$ time steps</i>, rather than every time step.
The constraint that $Q(s_t, a_t, s_{t+K}, K) = 0$ enforces the feasibility of
the trajectory. Visually, rather than explicitly planning $K$ steps and actions
like so 
</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/tdm/comp-mb.jpeg" alt="Model based planning many steps." width="80%"><br></p>

<p>
We can instead directly plan over $K$ time steps as shown below:
</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/tdm/comp-tdm.jpeg" alt="TDM planning one step" width="80%"><br></p>

<p>
As we increase $K$, we get temporally more and more abstract plans. In between
the $K$ time steps, we use a model-free approach to take actions, thereby
allowing the model-free policy &#8220;abstract away&#8221; the details of how the goal is
actually reached. For the biking problem and for large enough values of $K$, the
optimization could result in a plan like:
</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/tdm/tdm-plan-small.png" alt="A model-based planner can be used to choose temporally abstract goals. A model-free algorithm can be used to reach those goals." width="80%"><br><i>
A model-based planner can be used to choose temporally abstract goals. A model-free algorithm can be used to reach those goals.
</i>
</p>

<p>
One caveat is that this formulation can only optimize the reward at every $K$
steps. However, many tasks only care about some states, such as the final state
(e.g. &#8220;reach the Golden Gate Bridge&#8221;) and so this still captures a variety of
interesting tasks.
</p>

<h3>Related Work</h3>
<p>
  We&#8217;re not the first to look at the connection between model-based and model-free reinforcement. <a href="https://users.cs.duke.edu/~parr/icml08.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Parr &#8216;08</a> and <a href="https://pdfs.semanticscholar.org/61d4/897dbf7ced83a0eb830a8de0dd64abb58ebd.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Boyan &#8216;99</a>, are particularly related, though they focus mainly on tabular and linear function approximators. The idea of training a goal condition Q function was also explored in <a href="http://www.incompleteideas.net/papers/horde-aamas-11.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sutton &#8216;11</a> and <a href="http://proceedings.mlr.press/v37/schaul15.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Schaul &#8216;15</a>, in the context of robot navigation and Atari games. Lastly, the relabelling scheme that we use is inspired by the work of <a href="https://arxiv.org/abs/1707.01495" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Andrychowicz &#8216;17</a>.
</p>


<h2>Experiments</h2>
<p>
  We tested TDMs on five simulated continuous control tasks and one real-world robotics task. One of the simulated tasks is to train a robot arm to push a cylinder to a target position. An example of the final pushing TDM policy and the associate learning curves are shown below:
</p>

<table><tr><td>
      <img src="http://bair.berkeley.edu/static/blog/tdm/pusher-video-small.gif" alt="Pusher video" height="300"></td>
    <td>
      <img src="http://bair.berkeley.edu/static/blog/tdm/pusher-learning-curve.jpg" alt="Pusher learning curve" height="300"></td>
  </tr></table><p>
<i>
Left: TDM policy for reaching task. Right: Learning curves. TDM is blue (lower is better).
</i>
</p>

<p>
In the learning curve to the right, we plot the final distance to goal versus
the number of environment samples (lower is better). Our simulation controls the
robots at 20 Hz, meaning that 1000 steps corresponds to 50 seconds in the real
world. The dynamics of this environment are relatively easy to learn, meaning
that a model-based approach should excel. As expected, the model-based
approaches (purple curve) learns quickly--roughly 3000 steps, or 25 minutes--and
performs well. The TDM approach (blue curve) also learn quickly--roughly 2000
steps, or 17 minutes. The model-free DDPG (without TDMs) baseline eventually
solves the task, but requires many more training samples. One reason the TDM
approach learns so quickly is that it effective is a model-based methods in
disguise.
</p>

<p>
The story looks much better for model-free approaches when we move to locomotion
tasks, which have substantially harder dynamics. One of the locomotion tasks
involves training a quadruped robot to move to a certain position. The resulting
TDM policy is shown below on the left, along with the accompanying learning
curve on the right.
</p>

<p>
</p><table><tr><td>
      <img src="http://bair.berkeley.edu/static/blog/tdm/ant-video.gif" alt="Pusher video" height="300"></td>
    <td>
      <img src="http://bair.berkeley.edu/static/blog/tdm/ant-learning-curve.jpg" alt="Pusher learning curve" height="300"></td>
  </tr></table><p>
<i>
Left: TDM policy for locomotion task. Right: Learning curves. TDM is blue (lower is better).
</i>
</p>

<p>
  Just as we use trial-and-error rather than planning to master riding a bicycle, we expect model-free methods to perform better than model-based methods on these locomotion tasks. This is precisely what we see in the learning curve on the right: the model-based method plateaus in performance. The model-free DDPG method learns more slowly, but eventually outperforms the model-based approach. TDM manages to both learn quickly and achieve good final performance. There are more experiments in the paper, including training a real-world 7 degree-of-freedom Sawyer to reach positions. We encourage the readers to check them out!
</p>

<h2>Future Directions</h2>
<p>
  Temporal difference models provide a formalism and practical algorithm for interpolating from model-free to model-based control. However, there&#8217;s a lot of future work to be done. For one, the derivation assumes that the environment and policies are deterministic. In practice, most environments are stochastic. Even if they were deterministic, there are compelling reasons to use a stochastic policy in practice (see <a href="http://bair.berkeley.edu/blog/2017/10/06/soft-q-learning/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this blog post</a> for one example). Extending TDMs to this setting would help move TDMs to more realistic environments. Another idea would be to combine TDMs with alternative model-based planning optimization algorithms that the ones we used in the paper.  Lastly, we&#8217;d like to apply TDMs to more challenging tasks with real-world robots, like locomotion, manipulation, and, of course, bicycling to the Golden Gate Bridge.
</p>
<p>
  This work will be presented at ICLR 2018. For more information about TDMs, check out the following links and come see us at our poster presentation at ICLR in Vancouver:
</p>
<ul><li><a href="https://arxiv.org/abs/1802.09081" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ArXiv Preprint</a></li>
<li><a href="https://github.com/vitchyr/rlkit" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Code</a></li>
</ul><p>
  Let us know if you have any questions or comments!
</p>

<p>
  $^\dagger$ We call it a temporal difference model because we train $Q$ with temporal difference learning and use $Q$ as a model.
</p>
<hr><p>I would like to thank Sergey Levine and Shane Gu for their valuable feedback when preparing this blog post.</p>

&#60;!--
<h2 id="references">References</h2>
<ul>
  <li>Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob
McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba. Hindsight experience replay. arXiv
preprint arXiv:1707.01495, 2017.</li>
  <li>Justin A Boyan. Least-squares temporal difference learning. In Proceedings of the 16th International
Conference on Machine Learning, pp. 49&#8211;56, 1999.</li>
  <li>Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa,
David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. arXiv
preprint arXiv:1509.02971, 2015.</li>
  <li>Ronald Parr, Lihong Li, Gavin Taylor, Christopher Painter-Wakefield, and Michael L Littman. An
analysis of linear models, linear value-function approximation, and feature selection for reinforcement
learning. In International Conference on Machine learning, 2008.</li>
  <li>Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver. Universal value function approximators.
In Proceedings of the 32nd International Conference on Machine Learning, pp. 1312&#8211;1320, 2015.</li>
  <li>Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White,
and Doina Precup. Horde: A scalable real-time architecture for learning knowledge from unsupervised
sensorimotor interaction. In The 10th International Conference on Autonomous Agents and
Multiagent Systems-Volume 2, pp. 761&#8211;768. International Foundation for Autonomous Agents
and Multiagent Systems, 2011.</li>
</ul>
--&#62;
]]></description>
										<content:encoded><![CDATA[<p><strong>By Vitchyr Pong</strong><br />
You’ve decided that you want to bike from your house by UC Berkeley to the Golden Gate Bridge. It’s a nice 20 mile ride, but there’s a problem: you’ve never ridden a bike before!<span id="more-101074"></span>To make matters worse, you are new to the Bay Area, and all you have is a good ol’ fashion map to guide you. How do you get started?  Let’s first figure out how to ride a bike. One strategy would be to do a lot of studying and planning. Read books on how to ride bicycles. Study physics and anatomy. Plan out all the different muscle movements that you’ll make in response to each perturbation. This approach is noble, but for anyone who’s ever learned to ride a bike, they know that this strategy is doomed to fail. There’s only one way to learn how to ride a bike: trial and error. Some tasks like riding a bike are just too complicated to plan out in your head.  Once you’ve learned how to ride your bike, how would you get to the Golden Gate Bridge? You could reuse your trial-and-error strategy. Take a few random turns and see if you end up at the Golden Gate Bridge. Unfortunately, this strategy would take a very, very long time. For this sort of problem, planning is a much faster strategy, and requires considerably less real-world experience and trial-and-error. In reinforcement learning terms, it is more <i>sample-efficient</i>.  <!-- 

<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/riding-bike-small.png" alt="Some skills you learn by trial and error." /> <i> Some skills you learn by trial and error. </i>  

<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/hitchhiker-small.png" alt="Other times, planning ahead is better." /> <i> Other times, planning ahead is better. </i>  --> </p>
<table class="col-2">
<tbody>
<tr>
<td style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/riding-bike-small.png" height="260" /></td>
<td style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/hitchhiker-small.png" height="260" /></td>
</tr>
</tbody>
</table>
<p style="text-align: center;"><i> Left: some skills you learn by trial and error. Right: other times, planning ahead is better. </i></p>
<p> While simple, this thought experiment highlights some important aspects of human intelligence. For some tasks, we use a trial-and-error approach, and for others we use a planning approach. A similar phenomena seems to have emerged in reinforcement learning (RL). In the parlance of RL, empirical results show that some tasks are better suited for model-free (trial-and-error) approaches, and others are better suited for model-based (planning) approaches.  However, the biking analogy also highlights that the two systems are not completely independent. In particularly, to say that learning to ride a bike is <i>just</i> trial-and-error is an oversimplification. In fact, when learning to bike by trial-and-error, you’ll employ a bit of planning. Perhaps your plan will initially be, “Don’t fall over.” As you improve, you’ll make more ambitious plans, such as, “Bike forwards for two meters without falling over.” Eventually, your bike-riding skills will be so proficient that you can start to plan in very abstract terms (“Bike to the end of the road.”) to the point that all there is left to do is planning and you no longer need to worry about the nitty-gritty details of riding a bike. We see that there is a gradual transition from the model-free (trial-and-error) strategy to a model-based (planning) strategy. If we could develop artificial intelligence algorithms&#8211;and specifically RL algorithms&#8211;that mimic this behavior, it could result in an algorithm that both performs well (by using trial-and-error methods early on) and is sample efficient (by later switching to a planning approach to achieve more abstract goals).  This post covers temporal difference model (TDM), which is a RL algorithm that captures this smooth transition between model-free and model-based RL. Before describing TDMs, we start by first describing how a typical model-based RL algorithm works.  <!--more--> </p>
<h2 id="model-based-reinforcement-learning">Model-Based Reinforcement Learning</h2>
<p> In reinforcement learning, we have some some state space $\mathcal{S}$ and action space $\mathcal{A}$. If at time $t$ we are in state $s_t \in \mathcal{S}$ and take action $a_t\in \mathcal{A}$, we transition to a new state $s_{t+1} = f(s_t, a_t)$ according to a dynamics model $f: \mathcal{S} \times \mathcal{A} \mapsto \mathcal{S}$. The goal is to maximize rewards summed over the visited state: $\sum_{t=1}^{T-1} r(s_t, a_, s_{t+1})$. Model-based RL algorithms assume you are given (or learn) the dynamics model $f$. Given this dynamics model, there are a variety of model-based algorithms. For this post, we consider methods that perform the following optimization to choose a sequence of actions and states to maximize rewards:  <script type="math/tex; mode=display">   \qquad \text{max}_{a_{1:T-1}, s_{1:T}} \sum_{t=1}^{T-1} r(s_t, a_t, s_{t+1}) \text{ subject to }f(s_t, a_t) = s_{t+1} </script>  The optimization says to choose a sequence of states and actions that you maximize the rewards, while ensuring that the trajectory is feasible. Here, feasible means that each state-action-next-state transition is valid. For example, in the image below if you start in state $s_t$ and take action $a_t$, only the top $s_{t+1}$ results in a feasible transition. </p>
<p style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/bike-feasibility-compressed.png" alt="Planning a trip to the Golden Gate Bridge would be much easier if you could defy physics. However, the constraint in the model-based optimization problem ensures that only trajectories like the top row will be outputted. The bottom two trajectories may have high reward, but they’re not feasible." width="80%" /> <i> Planning a trip to the Golden Gate Bridge would be much easier if you could defy physics. However, the constraint in the model-based optimization problem ensures that only trajectories like the top row will be outputted. The bottom two trajectories may have high reward, but they’re not feasible. </i></p>
<p> In our biking problem, the optimization might result in a biking plan from Berkeley (top right) to the Golden Gate Bridge (middle left) that looks like this: </p>
<p style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/mb-bike-plan-small.png" alt="An example of a plan (states and actions) outputted the optimization problem." width="80%" /> <i> An example of a plan (states and actions) outputted the optimization problem. </i></p>
<p> While conceptually nice, this plan is not very realistic. Model-based approaches use a model $f(s, a)$ that predict the state at the very next time step. In robotics, a time step usually corresponds to a tenth or a hundredth of a second. So perhaps a more realistic depiction of the resulting plan might look like: </p>
<p style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/mb-short-time-small.png" alt="A more realistic plan." width="80%" /> <i> A more realistic plan. </i></p>
<p> If we think about how we plan in everyday life, we realize that we plan at much more temporally abstract terms. Rather than planning the position that our bike will be at the next tenth of a second, we make longer-term plan things like, “I will go to the end of the road.” Furthermore, we can only make these temporally abstract plans once we’ve learned how to ride a bike in the first place. As discussed earlier, we need some way to (1) start the learning using a trial-and-error approach and (2) provide a mechanism to gradually increase the level of abstraction that we use to plan. For this, we introduce temporal difference models. </p>
<h2 id="temporal-difference-models">Temporal Difference Models</h2>
<p> A temporal difference model (TDM)$^\dagger$, which we will write as $Q(s, a, s_g, \tau)$, is a function that, given a state $s \in \mathcal{S}$, action $a \in \mathcal{A}$, and goal state $s_g \in \mathcal{S}$, predicts how close an agent can get to the goal within $\tau$ time steps. Intuitively, a TDM answers the question, “If I try to bike to San Francisco in 30 minutes, how close will I get?” For robotics, a natural way to measure closeness is use Euclidean distance. </p>
<p style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/tdm-visualization-small.png" alt="A TDM predicts how close you will get to the goal (Golden Gate Bridge) after a fixed amount of time. After 30 minutes of biking, maybe you only reach the grey biker in the image above. In this case, the grey line represents the distance that the TDM should predict." width="80%" /> <i> A TDM predicts how close you will get to the goal (Golden Gate Bridge) after a fixed amount of time. After 30 minutes of biking, maybe you only reach the grey biker in the image above. In this case, the grey line represents the distance that the TDM should predict. </i></p>
<p> For those familiar with reinforcement learning, it turns out that a TDM can be viewed as a goal-conditioned Q function in a finite-horizon MDP. Because a TDM is just another Q function, we can train it with model-free (trial-and-error) algorithms. We use <a href="https://arxiv.org/abs/1509.02971" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">deep deterministic policy gradient</a> (DDPG) to train a TDM and retroactively relabel the goal and time horizon to increase the sample efficiency of our learning algorithm. In theory, any Q-learning algorithm could be used to train the TDM, but we found this to be effective. We encourage readers to check out the paper for more details. </p>
<h3 id="planning-with-a-tdm">Planning with a TDM</h3>
<p> Once we train a TDM, how can we use it to plan? It turns out that we can plan with the following optimization:  <script type="math/tex; mode=display">   \qquad \text{max}_{a_1, a_K, a_{2K}, s_1, s_K, s_{2K}, ..} \sum_{t=1, K, 2K, ...} r(s_t) \text{ subject to } Q(s_t, a_t, s_{t+K}, K) = 0 </script>  The intuition is similar to the model-based formulation. Choose a sequence of actions and states that maximize rewards and that are feasible. A key difference is that we only plan <i>every $K$ time steps</i>, rather than every time step. The constraint that $Q(s_t, a_t, s_{t+K}, K) = 0$ enforces the feasibility of the trajectory. Visually, rather than explicitly planning $K$ steps and actions like so </p>
<p style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/comp-mb.jpeg" alt="Model based planning many steps." width="80%" /></p>
<p> We can instead directly plan over $K$ time steps as shown below: </p>
<p style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/comp-tdm.jpeg" alt="TDM planning one step" width="80%" /></p>
<p> As we increase $K$, we get temporally more and more abstract plans. In between the $K$ time steps, we use a model-free approach to take actions, thereby allowing the model-free policy “abstract away” the details of how the goal is actually reached. For the biking problem and for large enough values of $K$, the optimization could result in a plan like: </p>
<p style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/tdm-plan-small.png" alt="A model-based planner can be used to choose temporally abstract goals. A model-free algorithm can be used to reach those goals." width="80%" /> <i> A model-based planner can be used to choose temporally abstract goals. A model-free algorithm can be used to reach those goals. </i></p>
<p> One caveat is that this formulation can only optimize the reward at every $K$ steps. However, many tasks only care about some states, such as the final state (e.g. “reach the Golden Gate Bridge”) and so this still captures a variety of interesting tasks. </p>
<h3 id="related-work">Related Work</h3>
<p> We’re not the first to look at the connection between model-based and model-free reinforcement. <a href="https://users.cs.duke.edu/~parr/icml08.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Parr ‘08</a> and <a href="https://pdfs.semanticscholar.org/61d4/897dbf7ced83a0eb830a8de0dd64abb58ebd.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Boyan ‘99</a>, are particularly related, though they focus mainly on tabular and linear function approximators. The idea of training a goal condition Q function was also explored in <a href="http://www.incompleteideas.net/papers/horde-aamas-11.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sutton ‘11</a> and <a href="http://proceedings.mlr.press/v37/schaul15.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Schaul ‘15</a>, in the context of robot navigation and Atari games. Lastly, the relabelling scheme that we use is inspired by the work of <a href="https://arxiv.org/abs/1707.01495" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Andrychowicz ‘17</a>. </p>
<h2 id="experiments">Experiments</h2>
<p> We tested TDMs on five simulated continuous control tasks and one real-world robotics task. One of the simulated tasks is to train a robot arm to push a cylinder to a target position. An example of the final pushing TDM policy and the associate learning curves are shown below: </p>
<table class="col-2">
<tbody>
<tr>
<td style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/pusher-video-small.gif" alt="Pusher video" height="300" /></td>
<td style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/pusher-learning-curve.jpg" alt="Pusher learning curve" height="300" /></td>
</tr>
</tbody>
</table>
<p style="text-align: center;"><i> Left: TDM policy for reaching task. Right: Learning curves. TDM is blue (lower is better). </i></p>
<p> In the learning curve to the right, we plot the final distance to goal versus the number of environment samples (lower is better). Our simulation controls the robots at 20 Hz, meaning that 1000 steps corresponds to 50 seconds in the real world. The dynamics of this environment are relatively easy to learn, meaning that a model-based approach should excel. As expected, the model-based approaches (purple curve) learns quickly&#8211;roughly 3000 steps, or 25 minutes&#8211;and performs well. The TDM approach (blue curve) also learn quickly&#8211;roughly 2000 steps, or 17 minutes. The model-free DDPG (without TDMs) baseline eventually solves the task, but requires many more training samples. One reason the TDM approach learns so quickly is that it effective is a model-based methods in disguise.  The story looks much better for model-free approaches when we move to locomotion tasks, which have substantially harder dynamics. One of the locomotion tasks involves training a quadruped robot to move to a certain position. The resulting TDM policy is shown below on the left, along with the accompanying learning curve on the right. </p>
<table class="col-2">
<tbody>
<tr>
<td style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/ant-video.gif" alt="Pusher video" height="300" /></td>
<td style="text-align: center;"><img decoding="async" src="http://bair.berkeley.edu/static/blog/tdm/ant-learning-curve.jpg" alt="Pusher learning curve" height="300" /></td>
</tr>
</tbody>
</table>
<p style="text-align: center;"><i> Left: TDM policy for locomotion task. Right: Learning curves. TDM is blue (lower is better). </i></p>
<p> Just as we use trial-and-error rather than planning to master riding a bicycle, we expect model-free methods to perform better than model-based methods on these locomotion tasks. This is precisely what we see in the learning curve on the right: the model-based method plateaus in performance. The model-free DDPG method learns more slowly, but eventually outperforms the model-based approach. TDM manages to both learn quickly and achieve good final performance. There are more experiments in the paper, including training a real-world 7 degree-of-freedom Sawyer to reach positions. We encourage the readers to check them out! </p>
<h2 id="future-directions">Future Directions</h2>
<p> Temporal difference models provide a formalism and practical algorithm for interpolating from model-free to model-based control. However, there’s a lot of future work to be done. For one, the derivation assumes that the environment and policies are deterministic. In practice, most environments are stochastic. Even if they were deterministic, there are compelling reasons to use a stochastic policy in practice (see <a href="http://bair.berkeley.edu/blog/2017/10/06/soft-q-learning/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this blog post</a> for one example). Extending TDMs to this setting would help move TDMs to more realistic environments. Another idea would be to combine TDMs with alternative model-based planning optimization algorithms that the ones we used in the paper. Lastly, we’d like to apply TDMs to more challenging tasks with real-world robots, like locomotion, manipulation, and, of course, bicycling to the Golden Gate Bridge.  This work will be presented at ICLR 2018. For more information about TDMs, check out the following links and come see us at our poster presentation at ICLR in Vancouver: </p>
<ul>
<li><a href="https://arxiv.org/abs/1802.09081" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ArXiv Preprint</a></li>
<li><a href="https://github.com/vitchyr/rlkit" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Code</a></li>
</ul>
<p> Let us know if you have any questions or comments!  $^\dagger$ We call it a temporal difference model because we train $Q$ with temporal difference learning and use $Q$ as a model.  </p>
<hr />
<p>  I would like to thank Sergey Levine and Shane Gu for their valuable feedback when preparing this blog post.  <!-- 

<h2 id="references">References</h2>

 

<ul>  	

<li>Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba. Hindsight experience replay. arXiv preprint arXiv:1707.01495, 2017.</li>

  	

<li>Justin A Boyan. Least-squares temporal difference learning. In Proceedings of the 16th International Conference on Machine Learning, pp. 49–56, 1999.</li>

  	

<li>Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971, 2015.</li>

  	

<li>Ronald Parr, Lihong Li, Gavin Taylor, Christopher Painter-Wakefield, and Michael L Littman. An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning. In International Conference on Machine learning, 2008.</li>

  	

<li>Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver. Universal value function approximators. In Proceedings of the 32nd International Conference on Machine Learning, pp. 1312–1320, 2015.</li>

  	

<li>Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup. Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction. In The 10th International Conference on Autonomous Agents and Multiagent Systems-Volume 2, pp. 761–768. International Foundation for Autonomous Agents and Multiagent Systems, 2011.</li>

 </ul>

 This article was initially published on the <a href="“http://bair.berkeley.edu/blog/2017/06/20/welcome/”" data-wpel-link="internal">BAIR blog</a>, and appears here with the authors’ permission.
--> </p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Shared autonomy via deep reinforcement learning</title>
		<link>https://robohub.org/shared-autonomy-via-deep-reinforcement-learning/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Tue, 24 Apr 2018 21:59:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://robohub.org/shared-autonomy-via-deep-reinforcement-learning/</guid>

					<description><![CDATA[
A blind, autonomous pilot (left), suboptimal human pilot (center), and combined human-machine team (right) play the Lunar Lander game.



Imagine a drone pilot remotely flying a quadrotor, using an onboard camera to navigate and land. Unfamiliar fl...]]></description>
										<content:encoded><![CDATA[<p><strong>By Siddharth Reddy</strong></p>
<p>Imagine a drone pilot remotely flying a quadrotor, using an onboard camera to navigate and land. Unfamiliar flight dynamics, terrain, and network latency can make this system challenging for a human to control. One approach to this problem is to train an autonomous agent to perform tasks like patrolling and mapping without human intervention. This strategy works well when the task is clearly specified and the agent can observe all the information it needs to succeed. Unfortunately, many real-world applications that involve human users do not satisfy these conditions: the user&#8217;s intent is often private information that the agent cannot directly access, and the task may be too complicated for the user to precisely define. For example, the pilot may want to track a set of moving objects (e.g., a herd of animals) and change object priorities on the fly (e.g., focus on individuals who unexpectedly appear injured). <i>Shared autonomy</i> addresses this problem by combining user input with automated assistance; in other words, <i>augmenting</i> human control instead of replacing it.</p>
<p><span id="more-100607"></span></p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/small-cat-gapped-opt.gif" /><br />
<br />
<i><br />
A blind, autonomous pilot (left), suboptimal human pilot (center), and combined human-machine team (right) play the Lunar Lander game.<br />
</i>
</p>
<h3>Background</h3>
<p>The idea of combining human and machine intelligence in a shared-control system goes back to the early days of Ray Goertz&#8217;s <a href="https://www.osti.gov/servlets/purl/1054625" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">master-slave manipulator</a> in 1949, Ralph Mosher&#8217;s <a href="http://www.dtic.mil/dtic/tr/fulltext/u2/701359.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Hardiman exoskeleton</a> in 1969, and Marvin Minsky&#8217;s call for <a href="https://spectrum.ieee.org/robotics/artificial-intelligence/telepresence-a-manifesto" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">telepresence</a> in 1980. After decades of research in robotics, human-computer interaction, and artificial intelligence, interfacing between a human operator and a remote-controlled robot remains a challenge. According to a <a href="https://www.cs.cmu.edu/~cga/drc/jfr-what.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">review</a> of the 2015 <a href="https://www.darpa.mil/program/darpa-robotics-challenge" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">DARPA Robotics Challenge</a>, &#8220;the most cost effective research area to improve robot performance is Human-Robot Interaction&#8230;.The biggest enemy of robot stability and performance in the DRC was operator errors. Developing ways to avoid and survive operator errors is crucial for real-world robotics. Human operators make mistakes under pressure, especially without extensive training and practice in realistic conditions.&#8221;</p>
<table class="col-3">
<tbody>
<tr>
<td style="text-align:center;">
      <img decoding="async" widt="100%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/goertz.png" />
    </td>
<td style="text-align:center;">
      <img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/bci.png" />
    </td>
<td style="text-align:center;">
      <img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/javdani.png" />
    </td>
</tr>
<td>
<p style="text-align:center;">
      <i>Master-slave robotic manipulator <a href="https://www.osti.gov/servlets/purl/1054625" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">(Goertz, 1949)</a></i>
    </p>
</td>
<td>
<p style="text-align:center;">
      <i>Brain-computer interface for neural prosthetics <a href="http://www.cell.com/neuron/abstract/S0896-6273(14)00739-9" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">(Shenoy &amp; Carmena, 2014)</a></i>
    </p>
</td>
<td>
<p style="text-align:center;">
      <i>Formalism for model-based shared autonomy <a href="https://arxiv.org/pdf/1503.07619.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">(Javdani et al., 2015)</a></i>
    </p>
</td>
</tbody>
</table>
<p>One research thrust in shared autonomy approaches this problem by inferring the user&#8217;s goals and autonomously acting to achieve them. Chapter 5 of Shervin Javdani&#8217;s <a href="https://repository.cmu.edu/cgi/viewcontent.cgi?article=2100&amp;context=dissertations" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Ph.D. thesis</a> contains an excellent review of the literature. Such methods have made progress toward better <a href="https://people.csail.mit.edu/jalonsom/docs/17-schwartig-autonomy-icra.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">driver assist</a>, <a href="https://arxiv.org/pdf/1503.05451.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">brain-computer interfaces for prosthetic limbs</a>, and <a href="http://www.roboticsproceedings.org/rss08/p16.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">assistive teleoperation</a>, but tend to require prior knowledge about the world; specifically, (1) a dynamics model that predicts the consequences of taking a given action in a given state of the environment, (2) the set of possible goals for the user, and (3) an observation model that describes the user&#8217;s behavior given their goal. Model-based shared autonomy algorithms are well-suited to domains in which this knowledge can be directly hard-coded or learned, but are challenged by unstructured environments with ill-defined goals and unpredictable user behavior. We approached this problem from a different angle, using <i>deep reinforcement learning</i> to implement <i>model-free</i> shared autonomy.</p>
<p><a href="https://arxiv.org/pdf/1708.05866.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Deep reinforcement learning</a> uses neural network function approximation to tackle the curse of dimensionality in high-dimensional, continuous state and action spaces, and has recently achieved remarkable success in training autonomous agents from scratch to <a href="https://www.nature.com/articles/nature14236" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">play video games</a>, <a href="https://www.nature.com/articles/nature24270" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">defeat human world champions at Go</a>, and <a href="https://arxiv.org/pdf/1504.00702.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">control robots</a>. We have taken preliminary steps toward answering the following question: <i>can deep reinforcement learning be useful for building flexible and practical assistive systems?</i></p>
<h3>Model-Free RL with a Human in the Loop</h3>
<p>To enable shared-control teleoperation with minimal prior assumptions, we devised a model-free deep reinforcement learning algorithm for shared autonomy. The key idea is to learn an end-to-end mapping from environmental observation and user input to agent action, with task reward as the only form of supervision. From the agent&#8217;s perspective, the user acts like a prior policy that can be fine-tuned, and an additional sensor generating observations from which the agent can implicitly decode the user&#8217;s private information. From the user&#8217;s perspective, the agent behaves like an adaptive interface that learns a personalized mapping from user commands to actions that maximizes task reward.</p>
<p style="text-align:center;">
  <img decoding="async" align="middle" width="75%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/deepassist-diagram.png" /></p>
<p><i><br />
Fig. 1: An overview of our human-in-the-loop deep Q-learning algorithm for model-free shared autonomy<br />
</i>
</p>
<p>One of the core challenges in this work was adapting standard deep RL techniques to leverage control input from a human without significantly interfering with the user&#8217;s feedback control loop or tiring them with a long training period. To address these issues, we used deep Q-learning to learn an approximate state-action value function that computes the expected future return of an action given the current environmental observation and the user&#8217;s input. Equipped with this value function, the assistive agent executes the closest high-value action to the user&#8217;s control input. The reward function for the agent is a combination of known terms computed for every state, and a terminal reward provided by the user upon succeeding or failing at the task. See Fig. 1 for a high-level schematic of this process.</p>
<h2>Learning to Assist</h2>
<p>
<a href="https://arxiv.org/pdf/1706.00155.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Prior work</a> has formalized shared autonomy as a partially-observable Markov decision process (POMDP) in which the user&#8217;s goal is initially unknown to the agent and must be inferred in order to complete the task. Existing methods tend to assume the following components of the POMDP are known ex-ante: (1) the dynamics of the environment, or the state transition distribution $T$; (2) the set of possible goals for the user, or the goal space $\mathcal{G}$; and (3) the user&#8217;s control policy given their goal, or the user model $\pi_h$. In our work, we relaxed these three standard assumptions. We introduced a model-free deep reinforcement learning method that is capable of providing assistance without access to this knowledge, but can also take advantage of a user model and goal space when they are known.
</p>
<p>
In our problem formulation, the transition distribution $T$, the user&#8217;s policy $\pi_h$, and the goal space $\mathcal{G}$ are no longer all necessarily known to the agent. The reward function, which depends on the user&#8217;s private information, is<br />
$$<br />
R(s, a, s&#8217;) = \underbrace{R_{\text{general}}(s, a, s&#8217;)}_\text{known} + \underbrace{R_{\text{feedback}}(s, a, s&#8217;)}_\text{unknown, but observed}.<br />
$$<br />
This decomposition follows a structure typically present in shared autonomy: there are some terms in the reward that are known, such as the need to avoid collisions. We capture these in $R_{\text{general}}$. $R_{\text{feedback}}$ is user-generated feedback that depends on their private information. We do not know this function. We merely assume the agent is informed when the user provides feedback (e.g., by pressing a button). In practice, the user might simply indicate once per trial whether the agent succeeded or not.
</p>
<h3>Incorporating User Input</h3>
<p>Our method jointly embeds the agent&#8217;s observation of the environment $s_t$ with the information from the user $u_t$ by simply concatenating them. Formally,<br />
$$<br />
\tilde{s}_t = \left[ \begin{array}{c} s_t \\ u_t \end{array} \right].<br />
$$<br />
The particular form of $u_t$ depends on the available information. When we do not know the set of possible goals $\mathcal{G}$ or the user&#8217;s policy given their goal $\pi_h$, as is the case for most of our experiments, we set $u_t$ to the user&#8217;s action $a^h_t$. When we know the goal space $\mathcal{G}$, we set $u_t$ to the inferred goal $\hat{g}_t$. In particular, for problems with known goal spaces and user models, we found that using <a href="https://www.aaai.org/Papers/AAAI/2008/AAAI08-227.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">maximum entropy inverse reinforcement learning</a> to infer $\hat{g}_t$ led to improved performance. For problems with known goal spaces but unknown user models, we found that under certain conditions we could improve performance by training an <a href="http://www.bioinf.jku.at/publications/older/2604.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">LSTM</a> recurrent neural network to predict $\hat{g}_t$ given the sequence of user inputs using a training set of rollouts produced by the unassisted user.
</p>
<h3>Q-Learning with User Control</h3>
<p>
Model-free reinforcement learning with a human in the loop poses two challenges: (1) maintaining informative user input and (2) minimizing the number of interactions with the environment. If the user input is a suggested control, consistently ignoring the suggestion and taking a different action can degrade the quality of user input, since humans rely on feedback from their actions to perform real-time control tasks. Popular on-policy algorithms like <a href="https://arxiv.org/abs/1502.05477" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">TRPO</a> are difficult to deploy in this setting since they give no guarantees on how often the user&#8217;s input is ignored. They also tend to require a large number of interactions with the environment, which is impractical for human users. Motivated by these two criteria, we turned to <a href="https://www.nature.com/articles/nature14236" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">deep Q-learning</a>.
</p>
<p>
<a href="http://www.gatsby.ucl.ac.uk/~dayan/papers/cjch.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Q-learning</a> is an off-policy algorithm, enabling us to address (1) by modifying the behavior policy used to select actions given their expected returns and the user&#8217;s input. Drawing inspiration from the minimal intervention principle embodied in recent work on <a href="https://people.csail.mit.edu/jalonsom/docs/17-schwartig-autonomy-icra.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">parallel autonomy</a> and <a href="https://cpb-us-e1.wpmucdn.com/sites.northwestern.edu/dist/5/1812/files/2017/08/17rss_broad-13p5s48.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">outer-loop stabilization</a>, we execute a feasible action closest to the user&#8217;s suggestion, where an action is feasible if it isn&#8217;t that much worse than the optimal action. Formally,<br />
$$<br />
\pi_{\alpha}(a \mid \tilde{s}, a^h) = \delta\left(a = \mathop{\arg\max}\limits_{\{a : Q'(\tilde{s}, a) \geq (1 &#8211; \alpha) Q'(\tilde{s}, a^\ast)\}} f(a, a^h)\right),<br />
$$<br />
where $f$ is an action-similarity function and $Q'(\tilde{s}, a) = Q(\tilde{s}, a) &#8211; \min_{a&#8217; \in \mathcal{A}} Q(\tilde{s}, a&#8217;)$ maintains a sane comparison for negative Q values. The constant $\alpha \in [0, 1]$ is a hyperparameter that controls the tolerance of the system to suboptimal human suggestions, or equivalently, the amount of assistance.
</p>
<p>
Mindful of (2), we note that off-policy Q-learning tends to be more sample-efficient than policy gradient and Monte Carlo value-based methods. The structure of our behavior policy also speeds up learning when the user is approximately optimal: for appropriately large $\alpha$, the agent learns to fine-tune the user&#8217;s policy instead of learning to perform the task from scratch. In practice, this means that during the early stages of learning, the combined human-machine team performs at least as well as the unassisted human instead of performing at the level of a random policy.
</p>
<h2>User Studies</h2>
<p>We applied our method to two real-time assistive control problems: the <a href="https://gym.openai.com/envs/LunarLander-v2/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Lunar Lander game</a> and a quadrotor landing task. Both tasks involved controlling motion using a discrete action space and low-dimensional state observations that include position, orientation, and velocity information. In both tasks, the human pilot had private information that was necessary to complete the task, but wasn&#8217;t capable of succeeding on their own.</p>
<h3>The Lunar Lander Game</h3>
<p>The objective of the game was to land the vehicle between the flags without crashing or flying out of bounds using two lateral thrusters and a main engine. The assistive copilot could observe the lander&#8217;s position, orientation, and velocity, but not the position of the flags.</p>
<table class="col-2">
<tbody>
<tr>
<td style="text-align:center;">
      <img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/human-solo-lander-opt.gif" />
    </td>
<td style="text-align:center;">
      <img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/human-assisted-lander-opt.gif" />
    </td>
</tr>
<td>
<p>
      <i><b>Human Pilot (Solo):</b> The human pilot can&#8217;t stabilize and keeps crashing.</i>
    </p>
</td>
<td>
<p>
      <i><b>Human Pilot + RL Copilot:</b> The copilot improves stability while giving the pilot enough freedom to land between the flags.</i>
    </p>
</td>
</tbody>
</table>
<p>Humans rarely beat the Lunar Lander game on their own, but with a copilot they did much better.</p>
<p style="text-align:center;">
<img decoding="async" width="49%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/lander-user-study-fig.png" title="Lunar Lander User Study Success vs. Crash Rates" /><br />
<br />
<i><br />
Fig. 2a: Success and crash rates averaged over 30 episodes.<br />
</i><br />
<br />
<img decoding="async" width="49%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/human-pilot-solo-traj.png" title="Lunar Lander User Study Solo Trajectories" /><br />
<img decoding="async" width="49%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/human-pilot-assisted-traj.png" title="Lunar Lander User Study Assisted Trajectories" /><br />
<br />
<i><br />
Fig. 2b-c: Trajectories followed by human pilots with and without a copilot on Lunar Lander. Red trajectories end in a crash or out of bounds, green in success, and gray in neither. The landing pad is marked by a star. For the sake of illustration, we only show data for a landing site on the left boundary.<br />
</i>
</p>
<p>
In simulation experiments with synthetic pilot models (not shown here), we also observed a significant benefit to explicitly inferring the goal (i.e., the location of the landing pad) instead of simply adding the user&#8217;s raw control input to the agent&#8217;s observations, suggesting that goal spaces and user models can and should be taken advantage of when they are available.
</p>
<p>One of the drawbacks of analyzing Lunar Lander is that the game interface<br />
and physics do not reflect the complexity and unpredictability of a real-world robotic shared autonomy task.<br />
To evaluate our method in a more realistic environment, we formulated a task for a human pilot flying a real quadrotor.</p>
<h3>Quadrotor Landing Task</h3>
<p>The objective of the task was to land a <a href="https://www.parrot.com/us/drones/parrot-ardrone-20-elite-edition#parrot-ardrone-20-elite-edition" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Parrot AR-Drone 2</a> on a small, square landing pad at some distance from its initial take-off position, such that the drone&#8217;s first-person camera was pointed at a random object in the environment (e.g., a red chair), without flying out of bounds or running out of time. The pilot used a keyboard to control velocity, and was blocked from getting a third-person view of the drone so that they had to rely on the drone&#8217;s first-person camera feed to navigate and land. The assistive copilot observed position, orientation, and velocity, but did not know which object the pilot wanted to look at.</p>
<table class="col-2">
<tbody>
<tr>
<td style="text-align:center;">
			<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/human-solo-quad-opt.gif" />
		</td>
<td style="text-align:center;">
      <img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/human-assisted-quad-opt.gif" />
		</td>
</tr>
<tr>
<td>
<p>
      <i><b>Human Pilot (Solo):</b> The pilot&#8217;s display only showed the drone&#8217;s first-person view, so pointing the camera was easy but finding the landing pad was hard.</i>
		</p>
</td>
<td>
<p>
      <i><b>Human Pilot + RL Copilot:</b> The copilot didn&#8217;t know where the pilot wanted to point the camera, but it knew where the landing pad was. Together, the pilot and copilot succeeded at the task.</i>
		</p>
</td>
</tr>
</tbody>
</table>
<p>Humans found it challenging to simultaneously point the camera at the desired scene and navigate to the precise location of a feasible landing pad under time constraints.<br />
The assistive copilot had little trouble navigating to and landing on the landing pad, but did not know where to point the camera because it did not know what the human wanted to observe after landing. Together, the human could focus on pointing the camera and the copilot could focus on landing precisely on the landing pad.</p>
<p style="text-align:center;">
<img decoding="async" width="49%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/quad-user-study-fig.png" title="Quadrotor User Study Success vs. Crash Rates" /><br />
<br />
<i><br />
Fig. 3a: Success and crash rates averaged over 20 episodes.<br />
</i><br />
<br />
<img decoding="async" width="49%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/human-pilot-solo-traj-quad.png" title="Quadrotor User Study Solo Trajectories" /><br />
<img decoding="async" width="49%" src="http://bair.berkeley.edu/static/blog/shared-autonomy/human-pilot-assisted-traj-quad.png" title="Quadrotor User Study Assisted Trajectories" /><br />
<br />
<i><br />
Fig. 3b-c: A bird&#8217;s-eye view of trajectories followed by human pilots with and without a copilot on the quadrotor landing task. Red trajectories end in a crash or out of bounds, green in success, and gray in neither. The landing pad is marked by a star.<br />
</i>
</p>
<p>Our results showed that combined pilot-copilot teams significantly outperform individual pilots and copilots.</p>
<h3>What&#8217;s Next?</h3>
<p>Our method has a major weakness: model-free deep reinforcement learning typically requires lots of training data, which can be burdensome for human users operating physical robots. We mitigated this issue in our experiments by pretraining the copilot in simulation without a human pilot in the loop. Unfortunately, this is not always feasible for real-world applications due to the difficulty of building high-fidelity simulators and designing rich user-agnostic reward functions $R_{\text{general}}$. We are currently exploring different approaches to this problem.</p>
<p></p>
<p>If you want to learn more, check out our pre-print on arXiv: <em>Siddharth Reddy, Anca Dragan, Sergey Levine, <a href="https://arxiv.org/abs/1802.01744" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Shared Autonomy via Deep Reinforcement Learning</a>, arXiv, 2018.</em></p>
<p>The paper will appear at <a href="http://www.roboticsconference.org/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Robotics: Science and Systems 2018</a> from June 26-30. To encourage replication and extensions, we have released <a href="https://github.com/rddy/deepassist" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our code</a>. Additional videos are available through the <a href="https://sites.google.com/view/deep-assist" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">project website</a>.</p>
<p> This article was initially published on the <a href="“http://bair.berkeley.edu/blog/2017/06/20/welcome/”" data-wpel-link="internal">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Towards a virtual stuntman</title>
		<link>https://robohub.org/towards-a-virtual-stuntman/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 11 Apr 2018 17:52:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://robohub.org/towards-a-virtual-stuntman/</guid>

					<description><![CDATA[<p>
<img width="750" src="http://bair.berkeley.edu/static/blog/stuntman/teaser.gif"><br><i>
Simulated humanoid performing a variety of highly dynamic and acrobatic skills.
</i>
</p>

<p>Motion control problems have become standard benchmarks for reinforcement
learning, and deep RL methods have been shown to be effective for a diverse
suite of tasks ranging from manipulation to locomotion. However, characters
trained with deep RL often exhibit unnatural behaviours, bearing artifacts such
as jittering, asymmetric gaits, and <a href="http://www.youtube.com/watch?v=hx_bgoTF7bs&#038;t=1m28s" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">excessive movement of
limbs</a>. Can we train our characters to produce more natural behaviours?</p>

<!--more-->

<p>A wealth of inspiration can be drawn from computer graphics, where the
physics-based simulation of natural movements have been a subject of intense
study for decades. The greater emphasis placed on motion quality is often
motivated by applications in film, visual effects, and games. Over the years, a
rich body of work in physics-based character animation have developed
controllers to produce robust and natural motions for a large corpus of <a href="https://www.youtube.com/watch?v=Mh8t_TuI3B4" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">tasks</a> and <a href="https://www.cs.ubc.ca/~van/papers/2011-TOG-quadruped/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">characters</a>.
These methods often leverage human insight to incorporate task-specific control
structures that provide strong inductive biases on the motions that can be
achieved by the characters (e.g. <a href="https://www.cs.ubc.ca/~van/papers/2013-TOG-MuscleBasedBipeds/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">finite-state
machines</a>, <a href="http://www.delasa.net/slip/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reduced
models</a>, and <a href="http://mrl.snu.ac.kr/research/ProjectManyMuscle/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">inverse
dynamics</a>). But as a result of these design decisions, the controllers are
often specific to a particular character or task, and controllers developed for
walking may not extend to more dynamic skills, where human insight becomes
scarce.</p>

<p>In this work, we will draw inspiration from the two fields to take advantage of
the generality afforded by deep learning models while also producing
naturalistic behaviours that rival the state-of-the-art in full body motion
simulation in computer graphics. We present a conceptually simple RL framework
that enables simulated characters to learn highly dynamic and acrobatic skills
from reference motion clips, which can be provided in the form of mocap data
recorded from human subjects. Given a single demonstration of a skill, such as a
spin-kick or a backflip, our character is able to learn a robust policy to
imitate the skill in simulation. Our policies produce motions that are nearly
indistinguishable from mocap.</p>

<div>
  
</div>

<h1>Motion Imitation</h1>

<p>In most RL benchmarks, simulated characters are represented using simple models
that provide only a crude approximation of real world dynamics. Characters are
therefore prone to exploiting idiosyncrasies of the simulation to develop
unnatural behaviours that are infeasible in the real world. Incorporating more
realistic <a href="https://www.crowdai.org/challenges/nips-2017-learning-to-run" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">biomechanical
models</a> can lead to more natural behaviours. But constructing high-fidelity
models can be extremely challenging, and the resulting motions may nonetheless
be unnatural.</p>

<p>An alternative is to take a data-driven approach, where reference motion capture
of humans provides examples of natural motions. The character can then be
trained to produce more natural behaviours by imitating the reference motions.
Imitating motion data in simulation has a 
<a href="http://graphics.cs.cmu.edu/?p=671" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">long</a> <a href="https://dl.acm.org/citation.cfm?id=2422388" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">history</a> in 
computer animation and has seen some recent 
<a href="https://xbpeng.github.io/projects/DeepLoco/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">demonstrations with deep RL</a>. 
While the results do appear more natural, they are still far from being able to
faithfully reproduce a wide variety of motions.</p>

<p>In this work, our policies will be trained through a motion imitation task,
where the goal of the character is to reproduce a given kinematic reference
motion. Each reference motion is represented by a sequence of target poses
${\hat{q}_0, \hat{q}_1,\ldots,\hat{q}_T}$, where $\hat{q}_t$ is the target
pose at timestep $t$. The reward function is to minimize the least squares pose
error between the target pose $\hat{q}_t$ and the pose of the simulated
character $q_t$,</p>

<p>While more sophisticated methods have been applied for motion imitation, we
found that simply minimizing the tracking error (along with a couple of
additional insights) works surprisingly well. The policies are trained by
optimizing this objective using <a href="https://arxiv.org/abs/1707.06347" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PPO</a>.</p>

<p>With this framework, we are able to develop policies for a rich repertoire of
challenging skills ranging from locomotion to acrobatics, martial arts to
dancing.</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_sideflip.gif" height="160"><img src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_cartwheel.gif" height="160"><img src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_kipup.gif" height="160"><img src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_speed_vault.gif" height="160"><br><i>
The humanoid learns to imitate various skills. The blue character is the
simulated character, and the green character is replaying the respective mocap
clip. Top left: sideflip. Top right: cartwheel. Bottom left: kip-up. Bottom
right: speed vault.
</i>
</p>

<p>Next, we compare our method with previous results that used (e.g. <a href="https://arxiv.org/abs/1707.02201" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">generative adversarial imitation
learning (GAIL)</a>) to imitate mocap clips. Our method is substantially simpler
than GAIL and it is able to better reproduce the reference motions. The
resulting policy avoids many of the artifacts commonly exhibited by deep RL
methods, and enables the character to produce a fluid life-like running gait.</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_run.gif" height="250"><img src="http://bair.berkeley.edu/static/blog/stuntman/deepmind.gif" height="250"><br><i>
Comparison of our method (left) and work from Merel et al. [2017] using GAIL to
imitate mocap data. Our motions appear significantly more natural than previous
work using deep RL.
</i>
</p>

<h1>Insights</h1>

<h2>Reference State Initialization (RSI)</h2>

<p>Suppose the character is trying to imitate a backflip. How would it know that
doing a full rotation midair will result in high rewards? Since most RL
algorithms are retrospective, they only observe rewards for states they have
visited. In the case of a backflip, the character will have to observe
successful trajectories of a backflip before it learns that those states will
yield high rewards. But since a backflip can be very sensitive to the initial
conditions at takeoff and landing, the character is unlikely to accidentally
execute a successful trajectory through random exploration. To give the
character a hint, at the start of each episode, we will initialize the character
to a state sampled randomly along the reference motion. So sometimes the
character will start on the ground, and sometimes it will start in the middle of
the flip. This allows the character to learn which states will result in high
rewards even before it has acquired the proficiency to reach those states.</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/stuntman/no_rsi.png" height="280"><img src="http://bair.berkeley.edu/static/blog/stuntman/rsi.png" height="280"><br><i>
RSI provides the character with a richer initial state distribution by
initializing it to random point along the reference motion.
</i>
</p>

<p>Below is a comparison of the backflip policy trained with RSI and without RSI,
where the character is always initialized to a fixed initial state at the start
of the motion. Without RSI, instead of learning a flip, the policy just cheats
by hopping backwards.</p>

<p>
<img width="750" src="http://bair.berkeley.edu/static/blog/stuntman/backflip_ablation.gif"><br><i>
Comparison of policies trained without RSI or ET. RSI and ET can be crucial for
learning more dynamics motions. Left: RSI+ET. Middle: No RSI. Right: No ET.
</i>
</p>

<h2>Early Termination (ET)</h2>

<p>Early termination is a staple for RL practitioners, and it is often used to
improve simulation efficiency. If the character gets stuck in a state from which
there is no chance of success, then the episode is terminated early, to avoid
simulating the rest. Here we show that early termination can in fact have a
significant impact on the results. Again, let&#8217;s consider a backflip. During the
early stages of training, the policy is terrible and the character will spend
most of its time falling. Once the character has fallen, it can be extremely
difficult for it to recover. So the rollouts will be dominated by samples where
the character is just struggling in vain on the ground. This is analogous to the
class imbalance problem encountered by other methodologies such as supervised
learning. This issue can be mitigated by terminating an episode as soon as the
character enters such a futile state (e.g. falling). Coupled with RSI, ET helps
to ensure that a larger portion of the dataset consists of samples close to the
reference trajectory. Without ET the character never learns to perform a flip.
Instead, it just falls and then tries to mime the motion on the ground.</p>

<h1>More Results</h1>

<p>In total, we have been able to learn over 24 skills for the humanoid just by
providing it with different reference motions.</p>

<p>
<img width="750" src="http://bair.berkeley.edu/static/blog/stuntman/all_skills.gif"><br><i>
Humanoid trained to imitate a rich repertoire of skills.
</i>
</p>

<p>In addition to imitating mocap clips, we can also train the humanoid to perform
some additional tasks like kicking a randomly placed target, or throwing a ball
to a target.</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_strikc_spinkick.gif" height="225"><img src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_throw.gif" height="225"><br><i>
Policies trained to kick and throw a ball to a random target.
</i>
</p>

<p>We can also train a simulated Atlas robot to imitate mocap clips from a human.
Though the Atlas has a very different morphology and mass distribution, it is
still able to reproduce the desired motions. Not only can the policies imitate
the reference motions, they can also recover from pretty significant
perturbations.</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/stuntman/atlas_spinkick.gif" height="200"><img src="http://bair.berkeley.edu/static/blog/stuntman/atlas_backflip.gif" height="200"><br><i>
Atlas trained to perform a spin-kick and backflip. The policies are robust to
significant perturbations.
</i>
</p>

<p>But what do we do if we don&#8217;t have mocap clips? Suppose we want to simulate a
T-Rex. For various <a href="https://www.nationalgeographic.com/science/prehistoric-world/dinosaur-extinction/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reasons</a>,
it is a bit difficult to mocap a T-Rex. So instead, we can have an artist
hand-animate some keyframes and then train a policy to imitate those.</p>

<p>
<img width="750" src="http://bair.berkeley.edu/static/blog/stuntman/t-rex.gif"><br><i>
Simulated T-Rex trained to imitate artist-authored keyframes.
</i>
</p>

<p>By why stop at a T-Rex? Let&#8217;s train a lion:</p>

<p>
<img width="750" src="http://bair.berkeley.edu/static/blog/stuntman/lion3d_run.gif"><br><i>
Simulated lion. Reference motion courtesy of <a href="https://zivadynamics.com/lion-project" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Ziva Dynamics</a>.
</i>
</p>

<p>and a dragon:</p>

<p>
<img width="750" src="http://bair.berkeley.edu/static/blog/stuntman/dragon.gif"><br><i>
Simulated dragon with a 418D state space and 94D action space.
</i>
</p>

<p>The story here is that a simple method ends up working surprisingly well. Just
by minimizing the tracking error, we are able to train policies for a diverse
collection of characters and skills. We hope this work will help inspire the
development of more dynamic motor skills for both simulated characters and
robots in the real world. Exploring methods for imitating motions from more
prevalent sources such as video is also an exciting avenue for scenarios that
are challenging to mocap, such as animals and cluttered environments.</p>

<p>To learn more, <a href="https://xbpeng.github.io/projects/DeepMimic/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">check out our paper</a>.</p>

<p>We would like to thank the co-authors of this work: Pieter Abbeel, Sergey
Levine, and Michiel van de Panne. This project was done in collaboration with
the University of British Columbia.</p>]]></description>
										<content:encoded><![CDATA[<p>Motion control problems have become standard benchmarks for reinforcement learning, and deep RL methods have been shown to be effective for a diverse suite of tasks ranging from manipulation to locomotion. However, characters trained with deep RL often exhibit unnatural behaviours, bearing artifacts such as jittering, asymmetric gaits, and <a href="http://www.youtube.com/watch?v=hx_bgoTF7bs&amp;t=1m28s" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">excessive movement of limbs</a>. Can we train our characters to produce more natural behaviours?</p>
<p>  <span id="more-100263"></span> </p>
<p style="text-align:center;"> <img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/stuntman/teaser.gif" /> <br /> <i> Simulated humanoid performing a variety of highly dynamic and acrobatic skills. </i> </p>
<p>A wealth of inspiration can be drawn from computer graphics, where the physics-based simulation of natural movements have been a subject of intense study for decades. The greater emphasis placed on motion quality is often motivated by applications in film, visual effects, and games. Over the years, a rich body of work in physics-based character animation have developed controllers to produce robust and natural motions for a large corpus of <a href="https://www.youtube.com/watch?v=Mh8t_TuI3B4" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">tasks</a> and <a href="https://www.cs.ubc.ca/~van/papers/2011-TOG-quadruped/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">characters</a>. These methods often leverage human insight to incorporate task-specific control structures that provide strong inductive biases on the motions that can be achieved by the characters (e.g. <a href="https://www.cs.ubc.ca/~van/papers/2013-TOG-MuscleBasedBipeds/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">finite-state machines</a>, <a href="http://www.delasa.net/slip/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reduced models</a>, and <a href="http://mrl.snu.ac.kr/research/ProjectManyMuscle/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">inverse dynamics</a>). But as a result of these design decisions, the controllers are often specific to a particular character or task, and controllers developed for walking may not extend to more dynamic skills, where human insight becomes scarce.</p>
<p>In this work, we will draw inspiration from the two fields to take advantage of the generality afforded by deep learning models while also producing naturalistic behaviours that rival the state-of-the-art in full body motion simulation in computer graphics. We present a conceptually simple RL framework that enables simulated characters to learn highly dynamic and acrobatic skills from reference motion clips, which can be provided in the form of mocap data recorded from human subjects. Given a single demonstration of a skill, such as a spin-kick or a backflip, our character is able to learn a robust policy to imitate the skill in simulation. Our policies produce motions that are nearly indistinguishable from mocap.</p>
<div class="keep-aspect"><iframe title="SIGGRAPH 2018: DeepMimic paper (main video)" width="500" height="281" src="https://www.youtube-nocookie.com/embed/vppFvq2quQ0?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></div>
<p></p>
<h1 id="motion-imitation">Motion Imitation</h1>
<p>In most RL benchmarks, simulated characters are represented using simple models that provide only a crude approximation of real world dynamics. Characters are therefore prone to exploiting idiosyncrasies of the simulation to develop unnatural behaviours that are infeasible in the real world. Incorporating more realistic <a href="https://www.crowdai.org/challenges/nips-2017-learning-to-run" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">biomechanical models</a> can lead to more natural behaviours. But constructing high-fidelity models can be extremely challenging, and the resulting motions may nonetheless be unnatural.</p>
<p>An alternative is to take a data-driven approach, where reference motion capture of humans provides examples of natural motions. The character can then be trained to produce more natural behaviours by imitating the reference motions. Imitating motion data in simulation has a  <a href="http://graphics.cs.cmu.edu/?p=671" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">long</a> <a href="https://dl.acm.org/citation.cfm?id=2422388" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">history</a> in  computer animation and has seen some recent  <a href="https://xbpeng.github.io/projects/DeepLoco/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">demonstrations with deep RL</a>.  While the results do appear more natural, they are still far from being able to faithfully reproduce a wide variety of motions.</p>
<p>In this work, our policies will be trained through a motion imitation task, where the goal of the character is to reproduce a given kinematic reference motion. Each reference motion is represented by a sequence of target poses ${\hat{q}_0, \hat{q}_1,\ldots,\hat{q}_T}$, where $\hat{q}_t$ is the target pose at timestep $t$. The reward function is to minimize the least squares pose error between the target pose $\hat{q}_t$ and the pose of the simulated character $q_t$,</p>
<p>  <script type="math/tex; mode=display">r_t = {\rm exp}\Big[-2 \|\hat{q}_t - q_t \|^2 \Big]</script>  </p>
<p>While more sophisticated methods have been applied for motion imitation, we found that simply minimizing the tracking error (along with a couple of additional insights) works surprisingly well. The policies are trained by optimizing this objective using <a href="https://arxiv.org/abs/1707.06347" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">PPO</a>.</p>
<p>With this framework, we are able to develop policies for a rich repertoire of challenging skills ranging from locomotion to acrobatics, martial arts to dancing.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_sideflip.gif" height="160" style="margin: 10px;" /> <img decoding="async" src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_cartwheel.gif" height="160" style="margin: 10px;" /> <img decoding="async" src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_kipup.gif" height="160" style="margin: 10px;" /> <img decoding="async" src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_speed_vault.gif" height="160" style="margin: 10px;" /> <br /> <i> The humanoid learns to imitate various skills. The blue character is the simulated character, and the green character is replaying the respective mocap clip. Top left: sideflip. Top right: cartwheel. Bottom left: kip-up. Bottom right: speed vault. </i> </p>
<p>Next, we compare our method with previous results that used (e.g. <a href="https://arxiv.org/abs/1707.02201" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">generative adversarial imitation learning (GAIL)</a>) to imitate mocap clips. Our method is substantially simpler than GAIL and it is able to better reproduce the reference motions. The resulting policy avoids many of the artifacts commonly exhibited by deep RL methods, and enables the character to produce a fluid life-like running gait.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_run.gif" height="250" style="margin: 10px;" /> <img decoding="async" src="http://bair.berkeley.edu/static/blog/stuntman/deepmind.gif" height="250" style="margin: 10px;" /> <br /> <i> Comparison of our method (left) and work from Merel et al. [2017] using GAIL to imitate mocap data. Our motions appear significantly more natural than previous work using deep RL. </i> </p>
<h1 id="insights">Insights</h1>
<h2 id="reference-state-initialization-rsi">Reference State Initialization (RSI)</h2>
<p>Suppose the character is trying to imitate a backflip. How would it know that doing a full rotation midair will result in high rewards? Since most RL algorithms are retrospective, they only observe rewards for states they have visited. In the case of a backflip, the character will have to observe successful trajectories of a backflip before it learns that those states will yield high rewards. But since a backflip can be very sensitive to the initial conditions at takeoff and landing, the character is unlikely to accidentally execute a successful trajectory through random exploration. To give the character a hint, at the start of each episode, we will initialize the character to a state sampled randomly along the reference motion. So sometimes the character will start on the ground, and sometimes it will start in the middle of the flip. This allows the character to learn which states will result in high rewards even before it has acquired the proficiency to reach those states.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/stuntman/no_rsi.png" height="280" style="margin: 10px;" /> <img decoding="async" src="http://bair.berkeley.edu/static/blog/stuntman/rsi.png" height="280" style="margin: 10px;" /> <br /> <i> RSI provides the character with a richer initial state distribution by initializing it to random point along the reference motion. </i> </p>
<p>Below is a comparison of the backflip policy trained with RSI and without RSI, where the character is always initialized to a fixed initial state at the start of the motion. Without RSI, instead of learning a flip, the policy just cheats by hopping backwards.</p>
<p style="text-align:center;"> <img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/stuntman/backflip_ablation.gif" /> <br /> <i> Comparison of policies trained without RSI or ET. RSI and ET can be crucial for learning more dynamics motions. Left: RSI+ET. Middle: No RSI. Right: No ET. </i> </p>
<h2 id="early-termination-et">Early Termination (ET)</h2>
<p>Early termination is a staple for RL practitioners, and it is often used to improve simulation efficiency. If the character gets stuck in a state from which there is no chance of success, then the episode is terminated early, to avoid simulating the rest. Here we show that early termination can in fact have a significant impact on the results. Again, let’s consider a backflip. During the early stages of training, the policy is terrible and the character will spend most of its time falling. Once the character has fallen, it can be extremely difficult for it to recover. So the rollouts will be dominated by samples where the character is just struggling in vain on the ground. This is analogous to the class imbalance problem encountered by other methodologies such as supervised learning. This issue can be mitigated by terminating an episode as soon as the character enters such a futile state (e.g. falling). Coupled with RSI, ET helps to ensure that a larger portion of the dataset consists of samples close to the reference trajectory. Without ET the character never learns to perform a flip. Instead, it just falls and then tries to mime the motion on the ground.</p>
<h1 id="more-results">More Results</h1>
<p>In total, we have been able to learn over 24 skills for the humanoid just by providing it with different reference motions.</p>
<p style="text-align:center;"> <img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/stuntman/all_skills.gif" /> <br /> <i> Humanoid trained to imitate a rich repertoire of skills. </i> </p>
<p>In addition to imitating mocap clips, we can also train the humanoid to perform some additional tasks like kicking a randomly placed target, or throwing a ball to a target.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_strikc_spinkick.gif" height="225" style="margin: 10px;" /> <img decoding="async" src="http://bair.berkeley.edu/static/blog/stuntman/humanoid_throw.gif" height="225" style="margin: 10px;" /> <br /> <i> Policies trained to kick and throw a ball to a random target. </i> </p>
<p>We can also train a simulated Atlas robot to imitate mocap clips from a human. Though the Atlas has a very different morphology and mass distribution, it is still able to reproduce the desired motions. Not only can the policies imitate the reference motions, they can also recover from pretty significant perturbations.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/static/blog/stuntman/atlas_spinkick.gif" height="200" style="margin: 10px;" /> <img decoding="async" src="http://bair.berkeley.edu/static/blog/stuntman/atlas_backflip.gif" height="200" style="margin: 10px;" /> <br /> <i> Atlas trained to perform a spin-kick and backflip. The policies are robust to significant perturbations. </i> </p>
<p>But what do we do if we don’t have mocap clips? Suppose we want to simulate a T-Rex. For various <a href="https://www.nationalgeographic.com/science/prehistoric-world/dinosaur-extinction/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">reasons</a>, it is a bit difficult to mocap a T-Rex. So instead, we can have an artist hand-animate some keyframes and then train a policy to imitate those.</p>
<p style="text-align:center;"> <img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/stuntman/t-rex.gif" /> <br /> <i> Simulated T-Rex trained to imitate artist-authored keyframes. </i> </p>
<p>By why stop at a T-Rex? Let’s train a lion:</p>
<p style="text-align:center;"> <img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/stuntman/lion3d_run.gif" /> <br /> <i> Simulated lion. Reference motion courtesy of <a href="https://zivadynamics.com/lion-project" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Ziva Dynamics</a>. </i> </p>
<p>and a dragon:</p>
<p style="text-align:center;"> <img decoding="async" width="750" src="http://bair.berkeley.edu/static/blog/stuntman/dragon.gif" /> <br /> <i> Simulated dragon with a 418D state space and 94D action space. </i> </p>
<p>The story here is that a simple method ends up working surprisingly well. Just by minimizing the tracking error, we are able to train policies for a diverse collection of characters and skills. We hope this work will help inspire the development of more dynamic motor skills for both simulated characters and robots in the real world. Exploring methods for imitating motions from more prevalent sources such as video is also an exciting avenue for scenarios that are challenging to mocap, such as animals and cluttered environments.</p>
<p>To learn more, <a href="https://xbpeng.github.io/projects/DeepMimic/index.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">check out our paper</a>.</p>
<p>We would like to thank the co-authors of this work: Pieter Abbeel, Sergey Levine, and Michiel van de Panne. This project was done in collaboration with the University of British Columbia. This article was initially published on the <a href="“http://bair.berkeley.edu/blog/2017/06/20/welcome/”" data-wpel-link="internal">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Learning robot objectives from physical human interaction</title>
		<link>https://robohub.org/learning-robot-objectives-from-physical-human-interaction/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Thu, 08 Feb 2018 22:35:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://robohub.org/learning-robot-objectives-from-physical-human-interaction/</guid>

					<description><![CDATA[<p>Humans physically interact with each other every day &#8211; from grabbing someone&#8217;s hand when they are about to spill their drink, to giving your friend a nudge to steer them in the right direction, physical interaction is an intuitive way to convey information about personal preferences and how to perform a task correctly.</p>

<p>So why aren&#8217;t we physically interacting with current robots the way we do with each other? Seamless physical interaction between a human and a robot requires a lot: lightweight robot designs, reliable torque or force sensors, safe and reactive control schemes, the ability to predict the intentions of human collaborators, and more! Luckily, robotics has made many advances in the design of <a href="http://www.roboticgizmos.com/wp-content/uploads/2016/10/20/jaco2.gif" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">personal robots</a> specifically developed with humans in mind.</p>

<p>However, consider the example from the beginning where you grab your friend&#8217;s hand as they are about to spill their drink. Instead of your friend who is spilling, imagine it was a robot. Because state-of-the-art robot planning and control algorithms typically assume human physical interventions are disturbances, once you let go of the robot, it will resume its erroneous trajectory and continue spilling the drink. The key to this gap comes from how robots reason about physical interaction: instead of thinking about <em>why</em> the human physically intervened and replanning in accordance with what the human wants, most robots simply resume their original behavior after the interaction ends.</p>

<p>We argue that <strong>robots should treat physical human interaction as useful information about how they should be doing the task</strong>. We formalize reacting to physical interaction as an objective (or reward) learning problem and propose a solution that enables robots to change their behaviors <em>while they are performing a task</em> according to the information gained during these interactions.</p>

<!--more-->

<h2>Reasoning About Physical Interaction: Unknown Disturbance versus Intentional Information</h2>

<p>The field of <em>physical human-robot interaction</em> (pHRI) studies the design, control, and planning problems that arise from close physical interaction between a human and a robot in a shared workspace. Prior research in pHRI has developed safe and responsive control methods to react to a physical interaction that happens while the robot is performing a task. Proposed by <a href="http://summerschool.stiff-project.org/fileadmin/pdf/Hog1985.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Hogan et. al.</a>, impedance control is one of the most commonly used methods to move a robot along a desired trajectory when there are people in the workspace. With this control method, the <strong>robot acts like a spring</strong>: it allows the person to push it, but moves back to an original desired position after the human stops applying forces. While this strategy is very fast and enables the robot to safely adapt to the human&#8217;s forces, the robot does not leverage these interventions to update its understanding of the task. Left alone, the robot would continue to perform the task in the same way as it had planned before any human interactions.</p>

<p><img src="http://bair.berkeley.edu/static/blog/phri/impedance_control.gif" alt="impedance_control" width="50%"></p>

<p>Why is this the case? It boils down to what assumptions the robot makes about its knowledge of the task and the meaning of the forces it senses. Typically, a robot is given a notion of its task in the form of an <em>objective function</em>. This objective function encodes rewards for different aspects of the task like  &#8220;reach a goal at location X&#8221;  or  &#8220;move close to the table while staying far away from people&#8221;. The robot uses its objective function to produce a motion that best satisfies all the aspects of the task: for example, the robot would move toward goal X while choosing a path that is far from a human and close to the table. If the robot&#8217;s original objective function was correct, then any physical interaction is simply a disturbance from its correct path. Thus, the robot should allow the physical interaction to perturb it for safety purposes, but it will return to the original path it planned since it stubbornly believes it is correct.</p>

<p>In contrast, we argue that human interventions are often intentional and occur because the robot is doing something wrong. While the robot&#8217;s original behavior may have been optimal with respect to its pre-defined objective function, the fact that a human intervention was necessary implies that <strong>the original objective function was not quite right</strong>. Thus, physical human interactions are no longer disturbances but rather informative observations about what the robot&#8217;s true objective should be. With this in mind, we take inspiration from <a href="http://ai.stanford.edu/~ang/papers/icml00-irl.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">inverse reinforcement learning</a> (IRL), where the robot observes some behavior (e.g., being pushed away from the table) and tries to infer an unknown objective function (e.g., &#8220;stay farther away from the table&#8221;). Note that while many IRL methods focus on the robot doing better <em>the next time</em> it performs the task, we focus on the robot completing its <em>current</em> task correctly.</p>

<h2>Formalizing Reacting to pHRI</h2>

<p>With our insight on physical human-robot interactions, we can formalize pHRI as a dynamical system, where the robot is unsure about the correct objective function and the human&#8217;s interactions provide it with information. This formalism defines a broad class of pHRI algorithms, which includes existing methods such as impedance control, and enables us to derive a novel online learning method.</p>

<p>We will focus on two parts of the formalism: (1) the structure of the objective function and (2) the observation model that lets the robot reason about the objective given a human physical interaction. Let  be the robot&#8217;s state (e.g., position and velocity) and  be the robot&#8217;s action (e.g., the torque it applies to its joints). The human can physically interact with the robot by applying an external torque, called , and the robot moves to the next state via its dynamics, .</p>

<h3>The Robot Objective: Doing the Task Right with Minimal Human Interaction</h3>

<p>In pHRI, we want the robot to learn from the human, but at the same time we do not want to overburden the human with constant physical intervention. Hence, we can write down an objective for the robot that optimizes both completing the task and minimizing the amount of interaction required, ultimately trading off between the two.</p>

<p>Here,  encodes the task-related features (e.g., &#8220;distance to table&#8221;, &#8220;distance to human&#8221;, &#8220;distance to goal&#8221;) and  determines the relative weight of each of these features. In the function,  encapsulates the true objective &#8211; if the robot knew exactly how to weight all the aspects of its task, then it could compute how to perform the task optimally. However, this parameter is not known by the robot! Robots will not always know the right way to perform a task, and certainly not the human-preferred way.</p>

<h3>The Observation Model: Inferring the Right Objective from Human Interaction</h3>

<p>As we have argued, the robot should observe the human&#8217;s actions to infer the unknown task objective. To link the direct human forces that the robot measures with the objective function, the robot uses an <em>observation model</em>. Building on prior work in  <a href="https://www.aaai.org/Papers/AAAI/2008/AAAI08-227.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">maximum entropy IRL</a> as well as the Bolzmann distributions used in <a href="http://web.mit.edu/clbaker/www/papers/cogsci2007.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">cognitive science models</a> of human behavior, we model the human&#8217;s interventions as corrections which approximately maximize the robot&#8217;s expected reward at state  while taking action . This expected reward emcompasses the immediate and future rewards and is captured by the -value:</p>

<p>Intuitively, this model says that a human is more likely to choose a physical correction that, when combined with the robot&#8217;s action, leads to a desirable (i.e., high-reward) behavior.</p>

<h2>Learning from Physical Human-Robot Interactions in Real-Time</h2>

<p>Much like teaching another human, we expect that the robot will continuously learn while we interact with it. However, the learning framework that we have introduced requires that the robot solve a Partially Observable Markov Decision Process (POMDP); unfortunately, it is well known that solving POMDPs exactly is at best computationally expensive, and at worst intractable. Nonetheless, we can derive approximations from this formalism that can enable the robot to learn and act while humans are interacting.</p>

<p>To achieve such in-task learning, we make three approximations summarized below:</p>

<p><strong>1) Separate estimating the true objective from solving for the optimal control policy.</strong> This means at every timestep, the robot updates its belief over possible  values, and then re-plans an optimal control policy with the new distribution.</p>

<p><strong>2) Separate planning from control</strong>. Computing an optimal control policy means computing the optimal action to take at every state in a continuous state, action, and belief space. Although re-computing a full optimal <em>policy</em> after every interaction is not tractable in real-time, we can re-compute an optimal <em>trajectory</em> from the current state in real-time. This means that the robot first plans a trajectory that best satisfies the current estimate of the objective, and then uses an impedance controller to track this trajectory.  The use of impedance control here gives us the nice properties described earlier, where people can physically modify the robot&#8217;s state while still being safe during interaction.</p>

<p>Looking back at our estimation step, we will make a similar shift to trajectory space and modify our observation model to reflect this:</p>

<p>Now, our observation model depends only on the cumulative reward  along a trajectory, which is easily computed by summing up the reward at each timestep. With this approximation, when reasoning about the true objective, the robot only has to consider the likelihood of a human&#8217;s preferred trajectory, , given the current trajectory it is executing, .</p>

<p>But what is the human&#8217;s preferred trajectory, ? The robot only gets to directly measure the human&#8217;s force $u_H$. One way to infer what is the human&#8217;s preferred trajectory is by propagating the human&#8217;s force throughout the robot&#8217;s current trajectory, . Figure 1. builds up the trajectory deformation based on prior work from <a href="http://dylanlosey.com/wp-content/uploads/2016/07/TRO_2017.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Losey and O&#8217;Malley</a>, starting from the robot&#8217;s original trajectory, then the force application, and then the deformation to produce .</p>

<p>
<img width="50%" src="http://bair.berkeley.edu/static/blog/phri/deformation_process.png" alt="deformation_process"><br><i>
Fig 1. To infer the human&#8217;s prefered trajectory given the current planned trajectory, the robot first measures the human&#8217;s interaction force, $u_H$, and then smoothly deforms the waypoints near interaction point to get the human&#8217;s preferred trajectory, $\xi_H$.
</i>
</p>

<p><strong>3) Plan with maximum a posteriori (MAP) estimate of </strong>. Finally, because  is a continuous variable and potentially high-dimensional, and since our observation model is not Gaussian, rather than planning with the full belief over , we will plan only with the MAP estimate. We find that the MAP estimate under a 2nd order Taylor Series Expansion about the robot&#8217;s current trajectory with a Gaussian prior is equivalent to running online gradient descent:</p>

<p>At every timestep, the robot updates its estimate of  in the direction of the cumulative feature difference, , between its current optimal trajectory and the human&#8217;s preferred trajectory. In the Learning from Demonstration literature, this update rule is analogous to online <a href="https://www.ri.cmu.edu/pub_files/pub4/ratliff_nathan_2006_1/ratliff_nathan_2006_1.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Max Margin Planning</a>; it is also analogous to <a href="https://arxiv.org/pdf/1601.00741.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">coactive learning</a>, where the user modifies waypoints for the current task to teach a reward function for future tasks.</p>

<p>Ultimately, putting these three steps together leads us to an elegant approximate solution to the original POMDP. At every timestep, the robot plans a trajectory  and begins to move. The human can physically interact, enabling the robot to sense their force $u_H$. The robot uses the human&#8217;s force to deform its original trajectory and produce the human&#8217;s desired trajectory, . Then the robot reasons about what aspects of the task are different between its original and the human&#8217;s preferred trajectory, and updates  in the direction of that difference. Using the new feature weights, the robot replans a trajectory that better aligns with the human&#8217;s preferences.</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/phri/algorithm.gif" alt="algorithm"></p>

<p>For a more thorough description of our formalism and approximations, please see <a href="http://proceedings.mlr.press/v78/bajcsy17a/bajcsy17a.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our recent paper from the 2017 Conference on Robot Learning</a>.</p>

<h2>Learning from Humans in the Real World</h2>

<p>To evaluate the benefits of in-task learning on a real personal robot, we recruited 10 participants for a user study. Each participant interacted with the robot running our proposed online learning method as well as a baseline where the robot did not learn from physical interaction and simply ran impedance control.</p>

<p>Fig 2. shows the three experimental household manipulation tasks, in each of which the robot started with an initially incorrect objective that participants had to correct. For example, the robot would move a cup from the shelf to the table, but without worrying about tilting the cup (perhaps not noticing that there is liquid inside).</p>

<p>
<img width="30%" src="http://bair.berkeley.edu/static/blog/phri/task1.png" title="cup"><img width="30%" src="http://bair.berkeley.edu/static/blog/phri/task2.png" title="table"><img width="30%" src="http://bair.berkeley.edu/static/blog/phri/task3.png" title="laptop"><br><i>
Fig 2. Trajectory generated with initial objective marked in black, and the desired trajectory from true objective in blue. Participants need to correct the robot to teach it to hold the cup upright (left), move closer to the table (center), and avoid going over the laptop (right).  </i>
</p>

<p>We measured the robot&#8217;s performance with respect to the true objective, the total effort the participant exerted, the total amount of interaction time, and the responses of a 7-point Likert scale survey.</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/phri/task1.gif" alt="cup gif"><br><i>
In Task 1, participants have to physically intervene when they see the robot tilting the cup and teach the robot to keep the cup upright.  
</i>
</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/phri/task2.gif" alt="table gif"><br><i>
Task 2 had participants teaching the robot to move closer to the table.
</i>
</p>

<p>
<img src="http://bair.berkeley.edu/static/blog/phri/task3.gif" alt="laptop gif"><br><i>
For Task 3, the robot&#8217;s original trajectory goes over a laptop. Participants have to physically teach the robot to move around the laptop instead of over it.
</i>
</p>

<p>The results of our user studies suggest that learning from physical interaction leads to better robot task performance with less human effort. Participants were able to <strong>get the robot to execute the correct behavior faster with less effort and interaction time</strong> when the robot was actively learning from their interactions during the task. Additionally, <strong>participants believed the robot understood their preferences more, took less effort to interact with, and was a more collaborative partner</strong>.</p>

<p>
<img width="30%" src="http://bair.berkeley.edu/static/blog/phri/taskCost_cameraready.png" title="task cost"><img width="30%" src="http://bair.berkeley.edu/static/blog/phri/taskEffort_cameraready.png" title="task effort"><img width="30%" src="http://bair.berkeley.edu/static/blog/phri/taskTime_cameraready.png" title="task time"><br><i>
Fig 3. Learning from interaction significantly outperformed not learning for each of our objective measures, including task cost, human effort, interaction time.
</i>
</p>

<p>Ultimately, we propose that robots should not treat human interactions as disturbances, but rather as informative actions. We showed that robots imbued with this sort of reasoning are capable of updating their understanding of the task they are performing and completing it correctly, rather than relying on people to guide them until the task is done.</p>

<p>This work is merely a step in exploring learning robot objectives from pHRI. Many open questions remain including developing solutions that can handle dynamical aspects (like preferences about the timing of the motion) and how and when to generalize learned objectives to new tasks. Additionally, robot reward functions will often have many task-related features and human interactions may only give information about a certain subset of relevant weights. Our recent work in HRI 2018 studied how a robot can disambiguate what the person is trying to correct by learning about only a single feature weight at a time. Overall, not only do we need algorithms that can learn from physical interaction with humans, but these methods must also reason about the inherent difficulties humans experience when trying to kinesthetically teach a complex &#8211; and possibly unfamiliar &#8211; robotic system.</p>

<hr><p>Thank you to Dylan Losey and Anca Dragan for their helpful feedback in writing this blog post.</p>

<hr><p>This post is based on the following papers:</p>

<ul><li>
    <p>A. Bajcsy* , D.P. Losey*, M.K. O&#8217;Malley, and A.D. Dragan. <strong>Learning Robot Objectives from Physical Human Robot Interaction</strong>. Conference on Robot Learning (CoRL), 2017.</p>
  </li>
  <li>
    <p>A. Bajcsy , D.P. Losey, M.K. O&#8217;Malley, and A.D. Dragan. <strong>Learning from Physical Human Corrections, One Feature at a Time</strong>. International Conference on Human-Robot Interaction (HRI), 2018.</p>
  </li>
</ul>]]></description>
										<content:encoded><![CDATA[<img decoding="async" src="http://bair.berkeley.edu/static/blog/phri/impedance_control.gif" alt="impedance_control" width="100%" />
<p>Humans physically interact with each other every day – from grabbing someone’s hand when they are about to spill their drink, to giving your friend a nudge to steer them in the right direction, physical interaction is an intuitive way to convey information about personal preferences and how to perform a task correctly.</p>
<p><span id="more-96855"></span></p>
<p>So why aren’t we physically interacting with current robots the way we do with each other? Seamless physical interaction between a human and a robot requires a lot: lightweight robot designs, reliable torque or force sensors, safe and reactive control schemes, the ability to predict the intentions of human collaborators, and more! Luckily, robotics has made many advances in the design of <a href="http://www.roboticgizmos.com/wp-content/uploads/2016/10/20/jaco2.gif" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">personal robots</a> specifically developed with humans in mind.</p>
<p>However, consider the example from the beginning where you grab your friend’s hand as they are about to spill their drink. Instead of your friend who is spilling, imagine it was a robot. Because state-of-the-art robot planning and control algorithms typically assume human physical interventions are disturbances, once you let go of the robot, it will resume its erroneous trajectory and continue spilling the drink. The key to this gap comes from how robots reason about physical interaction: instead of thinking about <em>why</em> the human physically intervened and replanning in accordance with what the human wants, most robots simply resume their original behavior after the interaction ends.</p>
<p>We argue that <strong>robots should treat physical human interaction as useful information about how they should be doing the task</strong>. We formalize reacting to physical interaction as an objective (or reward) learning problem and propose a solution that enables robots to change their behaviors <em>while they are performing a task</em> according to the information gained during these interactions.</p>
<h2 id="reasoning-about-physical-interaction-unknown-disturbance-versus-intentional-information">Reasoning About Physical Interaction: Unknown Disturbance versus Intentional Information</h2>
<p>The field of <em>physical human-robot interaction</em> (pHRI) studies the design, control, and planning problems that arise from close physical interaction between a human and a robot in a shared workspace. Prior research in pHRI has developed safe and responsive control methods to react to a physical interaction that happens while the robot is performing a task. Proposed by <a href="http://summerschool.stiff-project.org/fileadmin/pdf/Hog1985.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Hogan et. al.</a>, impedance control is one of the most commonly used methods to move a robot along a desired trajectory when there are people in the workspace. With this control method, the <strong>robot acts like a spring</strong>: it allows the person to push it, but moves back to an original desired position after the human stops applying forces. While this strategy is very fast and enables the robot to safely adapt to the human’s forces, the robot does not leverage these interventions to update its understanding of the task. Left alone, the robot would continue to perform the task in the same way as it had planned before any human interactions.</p>
<p>Why is this the case? It boils down to what assumptions the robot makes about its knowledge of the task and the meaning of the forces it senses. Typically, a robot is given a notion of its task in the form of an <em>objective function</em>. This objective function encodes rewards for different aspects of the task like  “reach a goal at location X”  or  “move close to the table while staying far away from people”. The robot uses its objective function to produce a motion that best satisfies all the aspects of the task: for example, the robot would move toward goal X while choosing a path that is far from a human and close to the table. If the robot’s original objective function was correct, then any physical interaction is simply a disturbance from its correct path. Thus, the robot should allow the physical interaction to perturb it for safety purposes, but it will return to the original path it planned since it stubbornly believes it is correct.</p>
<p>In contrast, we argue that human interventions are often intentional and occur because the robot is doing something wrong. While the robot’s original behavior may have been optimal with respect to its pre-defined objective function, the fact that a human intervention was necessary implies that <strong>the original objective function was not quite right</strong>. Thus, physical human interactions are no longer disturbances but rather informative observations about what the robot’s true objective should be. With this in mind, we take inspiration from <a href="http://ai.stanford.edu/~ang/papers/icml00-irl.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">inverse reinforcement learning</a> (IRL), where the robot observes some behavior (e.g., being pushed away from the table) and tries to infer an unknown objective function (e.g., “stay farther away from the table”). Note that while many IRL methods focus on the robot doing better <em>the next time</em> it performs the task, we focus on the robot completing its <em>current</em> task correctly.</p>
<h2 id="formalizing-reacting-to-phri">Formalizing Reacting to pHRI</h2>
<p>With our insight on physical human-robot interactions, we can formalize pHRI as a dynamical system, where the robot is unsure about the correct objective function and the human’s interactions provide it with information. This formalism defines a broad class of pHRI algorithms, which includes existing methods such as impedance control, and enables us to derive a novel online learning method.</p>
<p>We will focus on two parts of the formalism: (1) the structure of the objective function and (2) the observation model that lets the robot reason about the objective given a human physical interaction. Let <script type="math/tex">x</script> be the robot’s state (e.g., position and velocity) and <script type="math/tex">u_R</script> be the robot’s action (e.g., the torque it applies to its joints). The human can physically interact with the robot by applying an external torque, called <script type="math/tex">u_H</script>, and the robot moves to the next state via its dynamics, <script type="math/tex">\dot{x} = f(x,u_R+u_H)</script>.</p>
<h3 id="the-robot-objective-doing-the-task-right-with-minimal-human-interaction">The Robot Objective: Doing the Task Right with Minimal Human Interaction</h3>
<p>In pHRI, we want the robot to learn from the human, but at the same time we do not want to overburden the human with constant physical intervention. Hence, we can write down an objective for the robot that optimizes both completing the task and minimizing the amount of interaction required, ultimately trading off between the two.</p>
<p><script type="math/tex; mode=display">r(x,u_R,u_H;\theta) = \theta^{\top} \phi(x,u_R,u_H) - ||u_H||^2</script></p>
<p>Here, <script type="math/tex">\phi(x,u_R,u_H)</script> encodes the task-related features (e.g., “distance to table”, “distance to human”, “distance to goal”) and <script type="math/tex">\theta</script> determines the relative weight of each of these features. In the function, <script type="math/tex">\theta</script> encapsulates the true objective – if the robot knew exactly how to weight all the aspects of its task, then it could compute how to perform the task optimally. However, this parameter is not known by the robot! Robots will not always know the right way to perform a task, and certainly not the human-preferred way.</p>
<h3 id="the-observation-model-inferring-the-right-objective-from-human-interaction">The Observation Model: Inferring the Right Objective from Human Interaction</h3>
<p>As we have argued, the robot should observe the human’s actions to infer the unknown task objective. To link the direct human forces that the robot measures with the objective function, the robot uses an <em>observation model</em>. Building on prior work in  <a href="https://www.aaai.org/Papers/AAAI/2008/AAAI08-227.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">maximum entropy IRL</a> as well as the Bolzmann distributions used in <a href="http://web.mit.edu/clbaker/www/papers/cogsci2007.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">cognitive science models</a> of human behavior, we model the human’s interventions as corrections which approximately maximize the robot’s expected reward at state <script type="math/tex">x</script> while taking action <script type="math/tex">u_R+u_H</script>. This expected reward emcompasses the immediate and future rewards and is captured by the <script type="math/tex">Q</script>-value:</p>
<p><script type="math/tex; mode=display">P(u_H \mid x, u_R; \theta) \propto e^{Q(x,u_R+u_H;\theta)}</script></p>
<p>Intuitively, this model says that a human is more likely to choose a physical correction that, when combined with the robot’s action, leads to a desirable (i.e., high-reward) behavior.</p>
<h2 id="learning-from-physical-human-robot-interactions-in-real-time">Learning from Physical Human-Robot Interactions in Real-Time</h2>
<p>Much like teaching another human, we expect that the robot will continuously learn while we interact with it. However, the learning framework that we have introduced requires that the robot solve a Partially Observable Markov Decision Process (POMDP); unfortunately, it is well known that solving POMDPs exactly is at best computationally expensive, and at worst intractable. Nonetheless, we can derive approximations from this formalism that can enable the robot to learn and act while humans are interacting.</p>
<p>To achieve such in-task learning, we make three approximations summarized below:</p>
<p><strong>1) Separate estimating the true objective from solving for the optimal control policy.</strong> This means at every timestep, the robot updates its belief over possible <script type="math/tex">\theta</script> values, and then re-plans an optimal control policy with the new distribution.</p>
<p><strong>2) Separate planning from control</strong>. Computing an optimal control policy means computing the optimal action to take at every state in a continuous state, action, and belief space. Although re-computing a full optimal <em>policy</em> after every interaction is not tractable in real-time, we can re-compute an optimal <em>trajectory</em> from the current state in real-time. This means that the robot first plans a trajectory that best satisfies the current estimate of the objective, and then uses an impedance controller to track this trajectory.  The use of impedance control here gives us the nice properties described earlier, where people can physically modify the robot’s state while still being safe during interaction.</p>
<p>Looking back at our estimation step, we will make a similar shift to trajectory space and modify our observation model to reflect this:</p>
<p><script type="math/tex; mode=display">P(u_H \mid x, u_R; \theta) \propto e^{Q(x,u_R+u_H;\theta)} \rightarrow P(\xi_H \mid \xi_R; \theta) \propto e^{R(\xi_H, \xi_R;\theta)}</script></p>
<p>Now, our observation model depends only on the cumulative reward <script type="math/tex">R</script> along a trajectory, which is easily computed by summing up the reward at each timestep. With this approximation, when reasoning about the true objective, the robot only has to consider the likelihood of a human’s preferred trajectory, <script type="math/tex">\xi_H</script>, given the current trajectory it is executing, <script type="math/tex">\xi_R</script>.</p>
<p>But what is the human’s preferred trajectory, <script type="math/tex">\xi_H</script>? The robot only gets to directly measure the human’s force $u_H$. One way to infer what is the human’s preferred trajectory is by propagating the human’s force throughout the robot’s current trajectory, <script type="math/tex">\xi_R</script>. Figure 1. builds up the trajectory deformation based on prior work from <a href="http://dylanlosey.com/wp-content/uploads/2016/07/TRO_2017.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Losey and O’Malley</a>, starting from the robot’s original trajectory, then the force application, and then the deformation to produce <script type="math/tex">\xi_H</script>.</p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/phri/deformation_process.png" alt="deformation_process" /><br />
<i><br />
Fig 1. To infer the human’s prefered trajectory given the current planned trajectory, the robot first measures the human’s interaction force, $u_H$, and then smoothly deforms the waypoints near interaction point to get the human’s preferred trajectory, $\xi_H$.<br />
</i>
</p>
<p><strong>3) Plan with maximum a posteriori (MAP) estimate of <script type="math/tex">\theta</script></strong>. Finally, because <script type="math/tex">\theta</script> is a continuous variable and potentially high-dimensional, and since our observation model is not Gaussian, rather than planning with the full belief over <script type="math/tex">\theta</script>, we will plan only with the MAP estimate. We find that the MAP estimate under a 2nd order Taylor Series Expansion about the robot’s current trajectory with a Gaussian prior is equivalent to running online gradient descent:</p>
<p><script type="math/tex; mode=display">\theta^{t+1} = \theta^{t} + \alpha(\Phi(\xi^t_H) - \Phi(\xi^t_R))</script></p>
<p>At every timestep, the robot updates its estimate of <script type="math/tex">\theta</script> in the direction of the cumulative feature difference, <script type="math/tex">\Phi(\xi) = \sum_{x^t \in \xi} \phi(x^t)</script>, between its current optimal trajectory and the human’s preferred trajectory. In the Learning from Demonstration literature, this update rule is analogous to online <a href="https://www.ri.cmu.edu/pub_files/pub4/ratliff_nathan_2006_1/ratliff_nathan_2006_1.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Max Margin Planning</a>; it is also analogous to <a href="https://arxiv.org/pdf/1601.00741.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">coactive learning</a>, where the user modifies waypoints for the current task to teach a reward function for future tasks.</p>
<p>Ultimately, putting these three steps together leads us to an elegant approximate solution to the original POMDP. At every timestep, the robot plans a trajectory <script type="math/tex">\xi_R</script> and begins to move. The human can physically interact, enabling the robot to sense their force $u_H$. The robot uses the human’s force to deform its original trajectory and produce the human’s desired trajectory, <script type="math/tex">\xi_H</script>. Then the robot reasons about what aspects of the task are different between its original and the human’s preferred trajectory, and updates <script type="math/tex">\theta</script> in the direction of that difference. Using the new feature weights, the robot replans a trajectory that better aligns with the human’s preferences.</p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/phri/algorithm.gif" alt="algorithm" />
</p>
<p>For a more thorough description of our formalism and approximations, please see <a href="http://proceedings.mlr.press/v78/bajcsy17a/bajcsy17a.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our recent paper from the 2017 Conference on Robot Learning</a>.</p>
<h2 id="learning-from-humans-in-the-real-world">Learning from Humans in the Real World</h2>
<p>To evaluate the benefits of in-task learning on a real personal robot, we recruited 10 participants for a user study. Each participant interacted with the robot running our proposed online learning method as well as a baseline where the robot did not learn from physical interaction and simply ran impedance control.</p>
<p>Fig 2. shows the three experimental household manipulation tasks, in each of which the robot started with an initially incorrect objective that participants had to correct. For example, the robot would move a cup from the shelf to the table, but without worrying about tilting the cup (perhaps not noticing that there is liquid inside).</p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/phri/task1.png" title="cup" /><br />
<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/phri/task2.png" title="table" /><br />
<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/phri/task3.png" title="laptop" /><br />
<br />
<i><br />
Fig 2. Trajectory generated with initial objective marked in black, and the desired trajectory from true objective in blue. Participants need to correct the robot to teach it to hold the cup upright (left), move closer to the table (center), and avoid going over the laptop (right).  </i>
</p>
<p>We measured the robot’s performance with respect to the true objective, the total effort the participant exerted, the total amount of interaction time, and the responses of a 7-point Likert scale survey.</p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/phri/task1.gif" alt="cup gif" /><br />
<i><br />
In Task 1, participants have to physically intervene when they see the robot tilting the cup and teach the robot to keep the cup upright.<br />
</i>
</p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/phri/task2.gif" alt="table gif" /><br />
<i><br />
Task 2 had participants teaching the robot to move closer to the table.<br />
</i>
</p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/phri/task3.gif" alt="laptop gif" /><br />
<i><br />
For Task 3, the robot’s original trajectory goes over a laptop. Participants have to physically teach the robot to move around the laptop instead of over it.<br />
</i>
</p>
<p>The results of our user studies suggest that learning from physical interaction leads to better robot task performance with less human effort. Participants were able to <strong>get the robot to execute the correct behavior faster with less effort and interaction time</strong> when the robot was actively learning from their interactions during the task. Additionally, <strong>participants believed the robot understood their preferences more, took less effort to interact with, and was a more collaborative partner</strong>.</p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/phri/taskCost_cameraready.png" title="task cost" /><br />
<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/phri/taskEffort_cameraready.png" title="task effort" /><br />
<img decoding="async" width="100%" src="http://bair.berkeley.edu/static/blog/phri/taskTime_cameraready.png" title="task time" /><br />
<br />
<i><br />
Fig 3. Learning from interaction significantly outperformed not learning for each of our objective measures, including task cost, human effort, interaction time.<br />
</i>
</p>
<p>Ultimately, we propose that robots should not treat human interactions as disturbances, but rather as informative actions. We showed that robots imbued with this sort of reasoning are capable of updating their understanding of the task they are performing and completing it correctly, rather than relying on people to guide them until the task is done.</p>
<p>This work is merely a step in exploring learning robot objectives from pHRI. Many open questions remain including developing solutions that can handle dynamical aspects (like preferences about the timing of the motion) and how and when to generalize learned objectives to new tasks. Additionally, robot reward functions will often have many task-related features and human interactions may only give information about a certain subset of relevant weights. Our recent work in HRI 2018 studied how a robot can disambiguate what the person is trying to correct by learning about only a single feature weight at a time. Overall, not only do we need algorithms that can learn from physical interaction with humans, but these methods must also reason about the inherent difficulties humans experience when trying to kinesthetically teach a complex – and possibly unfamiliar – robotic system.</p>
<hr />
<p>Thank you to Dylan Losey and Anca Dragan for their helpful feedback in writing this blog post. This article was initially published on the <a href="“http://bair.berkeley.edu/blog/2017/06/20/welcome/”" data-wpel-link="internal">BAIR blog</a>, and appears here with the authors’ permission.</p>
<hr />
<p>This post is based on the following papers:</p>
<ul>
<li>
<p>A. Bajcsy* , D.P. Losey*, M.K. O’Malley, and A.D. Dragan. <strong>Learning Robot Objectives from Physical Human Robot Interaction</strong>. Conference on Robot Learning (CoRL), 2017.</p>
</li>
<li>
<p>A. Bajcsy , D.P. Losey, M.K. O’Malley, and A.D. Dragan. <strong>Learning from Physical Human Corrections, One Feature at a Time</strong>. International Conference on Human-Robot Interaction (HRI), 2018.</p>
</li>
</ul>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Physical adversarial examples against deep neural networks</title>
		<link>https://robohub.org/physical-adversarial-examples-against-deep-neural-networks/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sun, 31 Dec 2017 20:40:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<category><![CDATA[research]]></category>
		<guid isPermaLink="false">http://robohub.org/physical-adversarial-examples-against-deep-neural-networks/</guid>

					<description><![CDATA[This post is based on recent research by Ivan Evtimov, Kevin Eykholt, Earlence
Fernandes, Tadayoshi Kohno, Bo Li, Atul Prakash, Amir Rahmati, Dawn Song, and
Florian Tramèr.

Deep neural networks (DNNs) have enabled great progress in a variety of
applic...]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" src="http://robohub.org/wp-content/uploads/2017/12/CarStop.png" alt="" width="900" height="484" class="aligncenter size-full wp-image-94556" srcset="https://robohub.org/wp-content/uploads/2017/12/CarStop.png 900w, https://robohub.org/wp-content/uploads/2017/12/CarStop-425x229.png 425w, https://robohub.org/wp-content/uploads/2017/12/CarStop-768x413.png 768w" sizes="(max-width: 900px) 100vw, 900px" /><strong>By Ivan Evtimov, Kevin Eykholt, Earlence Fernandes, and Bo Li based on recent research by Ivan Evtimov, Kevin Eykholt, Earlence Fernandes, Tadayoshi Kohno, Bo Li, Atul Prakash, Amir Rahmati, Dawn Song, and Florian Tramèr.</strong>  </p>
<p>Deep neural networks (DNNs) have enabled great progress in a variety of application areas, including image processing, text analysis, and speech recognition. DNNs are also being incorporated as an important component in many cyber-physical systems. For instance, the vision system of a self-driving car can take advantage of DNNs to better recognize pedestrians, vehicles, and road signs. However, recent research has shown that DNNs are vulnerable to <em>adversarial examples</em>: Adding carefully crafted adversarial perturbations to the inputs can mislead the target DNN into mislabeling them during run time. Such adversarial examples raise security and safety concerns when applying DNNs in the real world. For example, adversarially perturbed inputs could mislead the perceptual systems of an autonomous vehicle into misclassifying road signs, with potentially catastrophic consequences.</p>
<p> <span id="more-94524"></span></p>
<p>There have been several techniques proposed to generate <em>adversarial examples</em> and to defend against them. In this blog post we will briefly introduce state-of-the-art algorithms to generate digital adversarial examples, and discuss our algorithm to generate <strong>physical</strong> adversarial examples on real objects under varying environmental conditions. We will also provide an update on our efforts to generate physical adversarial examples for object detectors.</p>
<p>  <!--more-->  </p>
<h1 id="digital-adversarial-examples">Digital Adversarial Examples</h1>
<p>Different methods have been proposed to generate adversarial examples in the white-box setting, where the adversary has full access to the DNN. The white-box setting assumes a powerful adversary and thus can help set the foundation for developing future fool-proof defenses. These methods contribute to understanding digital adversarial examples.</p>
<p>Goodfellow et al. proposed the <a href="https://arxiv.org/abs/1412.6572" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">fast gradient method</a> that applies a first-order approximation of the loss function to construct adversarial samples.</p>
<p><a href="https://arxiv.org/abs/1608.04644" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Optimization</a> based methods have also been proposed to create adversarial perturbations for targeted attacks.  Specifically, these attacks formulate an objective function whose solution seeks to maximize the difference between the true labeling of an input, and the attacker’s desired target labeling, while minimizing how different the inputs are, for some definition of input similarity.  In computer vision classification problems, a common measure is the L2-norm of the input vectors. Often, inputs with low L2 distances will be closer to each other. Thus, it is possible to compute inputs that are visually very similar to the human eye, but to a classifier, are very different.</p>
<p>Recent work has examined the <a href="https://arxiv.org/abs/1605.07277" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">black-box</a> transferability of digital adversarial examples, generating adversarial examples in black-box settings is also possible. These techniques involve generating adversarial examples for another known model in a white-box manner, and then running them against the target unknown model.</p>
<h1 id="physical-adversarial-examples">Physical Adversarial Examples</h1>
<p>To better understand these vulnerabilities, there has been extensive research on how <em>adversarial examples</em> may affect DNNs deployed in the <em>physical world</em>.</p>
<p><a href="https://arxiv.org/abs/1607.02533" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Kurakin et al.</a> showed that printed adversarial examples can be misclassified when viewed through a smartphone camera. <a href="https://www.cs.cmu.edu/~sbhagava/papers/face-rec-ccs16.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Sharif et al.</a> attacked face recognition systems by printing adversarial perturbations on the frames of eyeglasses. Their work demonstrated successful physical attacks in relatively stable physical conditions with little variation in pose, distance/angle from the camera, and lighting. This contributes an interesting understanding of physical examples in stable environments.</p>
<p>Our recent work “<a href="https://arxiv.org/abs/1707.08945" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Robust physical-world attacks on deep learning models</a>” has shown physical attacks on <strong>classifiers</strong>. (Check out the <a href="https://www.youtube.com/watch?v=1mJMPqi2bSQ&amp;feature=youtu.be" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">videos</a> <a href="https://www.youtube.com/watch?v=xwKpX-5Q98o&amp;feature=youtu.be" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">here</a>.) As the next logical step, we show attacks on object <strong>detectors</strong>. These computer vision algorithms identify relevant objects in a scene and predict bounding boxes indicating objects’ position and kind. Compared with classifiers, detectors are more challenging to fool as they process the entire image and can use contextual information (e.g. the orientation and position of the target object in the scene) in their predictions.</p>
<p>We demonstrate <em>physical</em> adversarial examples against the <a href="https://pjreddie.com/darknet/yolo/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">YOLO</a> detector, a popular state-of-the-art algorithm with good real-time performance. Our examples take the form of sticker perturbations that we apply to a real STOP sign. The following image shows our example physical adversarial perturbation.</p>
<p><img decoding="async" src="http://bair.berkeley.edu/static/blog/yolo/image1.png" width="1500" height="1999" class="alignnone size-full" /><br />
<img decoding="async" src="http://bair.berkeley.edu/static/blog/yolo/image4.png" width="1500" height="1999" class="alignnone size-full" /></p>
<p>We also perform dynamic tests by recording a video to test out the detection performance.  As can be seen in the video, the YOLO network does not perceive the STOP sign in almost all the frames. If a real autonomous vehicle were driving down the road with such an adversarial STOP sign, it would not see the STOP, possibly leading to a crash at an intersection. The perturbation we created is robust to changing distances and angles – the most commonly changing factors in a self-driving scenario.</p>
<p>More interestingly, the physical adversarial examples generated for the YOLO detector are also be able to fool standard <a href="https://github.com/endernewton/tf-faster-rcnn" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Faster-RCNN</a>. Our demo videos contains a dynamic test of the physical adversarial example on Faster-RCNN. As this is a black box attack on Faster-RCNN, the attack is not as successful as it is in the YOLO case. This is expected behavior. We believe that with additional techniques (such as ensemble training), the black box attack could be made more effective.  Additionally, specially optimizing an attack for Faster-RCNN will yield better results. We are currently working on a paper that explores these attacks in more detail. The image below is an example of Faster-RCNN not perceiving the STOP sign.</p>
<p><img decoding="async" src="http://bair.berkeley.edu/static/blog/yolo/image3.png" width="1503" height="1601" class="alignnone size-full" /><br />
<img decoding="async" src="http://bair.berkeley.edu/static/blog/yolo/image2.png" width="1853" height="1653" class="alignnone size-full" /></p>
<p>In both cases (YOLO and Faster-RCNN), a STOP sign is detected only when the camera is very close to the sign (about 3 to 4 feet away). In real settings, this distance is too close for a vehicle to take effective corrective action. Stay tuned for our upcoming paper that contains more details about the algorithm and results of physical perturbations against state-of-the-art object detectors.</p>
<h1 id="attack-algorithm-overview">Attack Algorithm Overview</h1>
<p>This algorithm is based off our earlier work on attacking classifiers. Fundamentally, we take an optimization approach to generating adversarial examples. However, our experimental experience indicates that generating robust physical adversarial examples for detectors requires simulating a larger set of varying physical conditions than what is needed to fool classifiers. This is likely because a detector takes much more contextual information into account while generating predictions. Key properties of the algorithm include the ability to specify sequences of physical condition simulations, and the ability to specify the translation invariance property. That is, a perturbation should be effective no matter where the target object is situated within the scene. As an object can move around freely in the scene depending on the viewer, perturbations not optimized for this property will likely break when the object moves. Our upcoming paper on this topic will contain more details on this algorithm.</p>
<h1 id="potential-defenses">Potential defenses</h1>
<p>Given these adversarial examples in both digital and physical world, potential defense methods have also been widely studied. Among them, different types of adversarial training methods are the most effective. <a href="https://arxiv.org/abs/1412.6572" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Goodfellow et al.</a> first proposed adversarial training as an effective way to improve the robustness of DNNs, and <a href="https://arxiv.org/abs/1705.07204" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Tramèr et al.</a> extend it to ensemble adversarial learning.  <a href="https://pdfs.semanticscholar.org/bcf1/1c7b9f4e155c0437958332507b0eaa44a12a.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Madry et al.</a> have also proposed robust networks via iterative training with adversarial examples. To conduct an adversarial training based defense, a large number of adversarial examples are required. In addition, these adversarial examples can make the defense more robust if they come from different models as suggested by work on <a href="https://arxiv.org/abs/1705.07204" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">ensemble training</a>. The benefit of ensemble adversarial training is to increase the diversity of adversarial examples so that the model can fully explore the adversarial example space. There are other types of defense methods as well, but <a href="http://nicholas.carlini.com/papers/2017_aisec_breakingdetection.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Carlini and Wagner</a> have shown that none of these existing defense method is robust enough given adaptive attack.</p>
<p>Overall, we are still a long way from finding the optimal defense strategy against these adversarial examples, and we are looking forward to exploring this exciting research area.</p>
<p><strong>Physical Adversarial Sticker Perturbations for YOLO </strong></p>
<div class="keep-aspect"><iframe title="Physical Adversarial Sticker Perturbations for YOLO" width="500" height="375" src="https://www.youtube-nocookie.com/embed/gkKyBmULVvM?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></div>
<p></p>
<p><strong>Physical Adversarial Examples for YOLO (2)</p>
<div class="keep-aspect"><iframe title="Physical Adversarial Examples for YOLO (2)" width="500" height="281" src="https://www.youtube-nocookie.com/embed/zSFZyzHdTO0?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></div>
<p></p>
<p><strong>Black box transfer to Faster RCNN of physical adversarial examples generated for YOLO</strong></p>
<div class="keep-aspect"><iframe title="Black box transfer to Faster RCNN of physical adversarial examples generated for YOLO" width="500" height="375" src="https://www.youtube-nocookie.com/embed/_ynduxh4uww?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></div>
<p></p>
<p>This article was initially published on the <a href="“http://bair.berkeley.edu/blog/2017/06/20/welcome/”" data-wpel-link="internal">BAIR blog</a>, and appears here with the authors’ permission.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Reverse curriculum generation for reinforcement learning agents</title>
		<link>https://robohub.org/reverse-curriculum-generation-for-reinforcement-learning-agents/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Sun, 24 Dec 2017 03:48:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://robohub.org/reverse-curriculum-generation-for-reinforcement-learning-agents/</guid>

					<description><![CDATA[<p>Reinforcement Learning (RL) is a powerful technique capable of solving complex tasks such as <a href="https://arxiv.org/abs/1506.02438" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">locomotion</a>, <a href="https://web.stanford.edu/class/psych209/Readings/MnihEtAlHassibis15NatureControlDeepRL.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Atari games</a>, <a href="https://arxiv.org/abs/1509.02971" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">racing games</a>, and <a href="https://arxiv.org/abs/1504.00702" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robotic manipulation tasks</a>, all through training an agent to optimize behaviors over a reward function. There are many tasks, however, for which it is <strong>hard to design a reward function that is both easy to train and that yields the desired behavior once optimized</strong>. Suppose we want a robotic arm to learn how to place a ring onto a peg. The most natural reward function would be for an agent to receive a reward of 1 at the desired end configuration and 0 everywhere else.  However, the required motion for this task&#8211;to align the ring at the top of the peg and then slide it to the bottom&#8211;is impractical to learn under such a binary reward, because the usual random exploration of our initial policy is unlikely to ever reach the goal, as seen in Video 1a. Alternatively, one can try to <a href="https://people.eecs.berkeley.edu/~pabbeel/cs287-fa09/readings/NgHaradaRussell-shaping-ICML1999.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">shape the reward function</a> to potentially alleviate this problem, but finding a good shaping <a href="https://arxiv.org/abs/1704.03073" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">requires considerable expertise and experimentation</a>. For example, directly minimizing the distance between the center of the ring and the bottom of the peg leads to an unsuccessful policy that smashes the ring against the peg, as in Video 1b. We propose a method to learn efficiently without modifying the reward function, by automatically generating a curriculum over start positions.</p>

<table><tr><td>
			<img src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_fail_cross.gif" alt="ring_fail_cross"></td>
    <td>
			<img src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_shapping_cross.gif" alt="ring_shapping_cross"></td>
  </tr><tr><td><p>
      <i>Video 1a: A randomly initialized policy is unable to reach the goal from most start positions, hence being unable to learn.</i>
		</p></td>
    <td><p>
      <i>Video 1b: Shaping the reward with a penalty on the distance from the ring center to the peg bottom yields an undesired behavior.
</i>
		</p></td>
  </tr></table><!--more--><h2>Curriculum instead of Reward Shaping</h2>

<p>We would like to train an agent to reach the goal from any starting position, without requiring an expert to shape the reward. Clearly, not all starting positions are equally difficult. In particular, even a random agent that is placed near to the goal will be able to reach the goal some of the time, receive a reward, and hence start learning! This acquired knowledge can then be bootstrapped to solve the task starting from further away from the goal. By <strong>choosing the ordering of the starting positions that we use in training</strong>, we can exploit this underlying structure of the problem and improve learning efficiency. A key advantage of this technique is that the <strong>reward function is not modified</strong>, and optimizing the sparse reward directly is less prone to yielding undesired behaviors. Ordering a set of related tasks to be learned is referred to as <strong>curriculum learning</strong>, and a central question for us is how to choose this task ordering. Our method, which we explain in more detail below, uses the performance of the learning agent to automatically generate a curriculum of tasks which start from the goal and expand outwards.</p>

<h3>Reverse Curriculum Intuition</h3>
<p>In <em>goal-oriented</em> tasks the aim is to reach a desired configuration from any start state. For example, in the ring-on-peg task introduced above, we desire to place the ring on the peg starting from any configuration. From most start positions, the random exploration of our initial policy never reaches the goal and hence perceives no reward. Nevertheless, it can be seen in Video 2a how a random policy is likely to reach the bottom of the peg if it is initialized from a nearby position. Then, once we have learned how to reach the goal from around the goal, learning from further away is easier since the agent already knows how to proceed if exploratory actions drive its state nearby the goal, as in Video 2b. Eventually, the agent successfully learns to reach  the goal from a wide range of starting positions, as in Video 2c.</p>

<p>
<img height="180" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_curr1.gif" title="ring_curr1"><img height="180" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_curr2.gif" title="ring_curr2"><img height="200" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_curr3.gif" title="ring_curr3"><br><i>
Video 2a-c: Our method strives to learn first starting nearby the goal, and then progressively expands in reverse the positions from where it starts.
</i>
</p>

<p>This method of learning in reverse, or growing outwards from the goal, draws inspiration from Dynamic Programming methods, where the solutions to easier sub-problems are used to compute the solution to harder problems.</p>

<h3>Starts of Intermediate Difficulty (SoID)</h3>

<p>To implement this reverse curriculum, we need to ensure that this outwards expansion happens at the right pace for the learning agent. In other words, we want to mathematically describe a set of starts that tracks the current agent performance and provides a good learning signal to our Reinforcement Learning algorithm. In particular we focus on Policy Gradient algorithms, which improve a parameterized policy by taking steps in the direction of an estimated gradient of the total expected reward . This gradient estimation is usually a variation of the original <a href="https://link.springer.com/article/10.1007/BF00992696" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">REINFORCE</a>, which is estimated by collecting $N$ on-policy trajectories  starting from states .</p>

<p>In <em>goal-oriented</em> tasks, the trajectory reward $R(\tau^i, s^i_0)$ is binary, indicating whether the agent reached the goal. Therefore, the usual baseline $R(\pi_i, s^i_0)$ estimates the probability of reaching the goal if the current policy $\pi_\theta$ is executed starting from $s^i_0$. Hence, we see from Eq. (1) that the terms of the sum corresponding to trajectories collected from starts $s^i_0$ that have success probability 0 or 1 will vanish. These are &#8220;wasted&#8221; trajectories as they do not contribute to the estimation of the gradient &#8211; they are either too hard or too easy. A similar analysis was already introduced in our prior <a href="https://arxiv.org/abs/1705.06366" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">work on multi-task RL</a>. In this case, to avoid training from starts from where our current policy either never gets to the goal or already masters it, we introduce the concept of &#8220;Start of Intermediate Difficulty&#8221; (SoID), which are start states $s_0$ that satisfy:</p>

<p>The values of  and  have the straightforward interpretation of minimum success probability acceptable for training from that start and maximum success probability above which we prefer to focus on training from other starts. In all our experiments we used 10% and 90%.</p>

<h2>Automatic Generation of the Reverse Curriculum</h2>

<p>From the above intuition and derivation, we would like to train our policy with trajectories starting from SoID states. Unfortunately, finding all starts that exactly satisfy Eq. (2) at every policy update is intractable, and hence we introduce an efficient approximation to automatically generate this reverse curriculum: we sample states nearby the starts that were estimated to be SoID during the previous iteration. To do that, we propose a way to <em>filter out non-SoID starts using the trajectories collected during the last training iteration and then sample nearby states</em>. The full algorithm is illustrated in Video 3 and details are given below.</p>

<div>
  
</div>

<p>
<i>
Video 3: Animation illustrating the main steps of our algorithm, and how it automatically produces a curriculum adapted to the current agent performance.
</i>
</p>

<h3>Filtering out Non-SoID</h3>

<p>At every policy gradient training iteration, we collect  trajectories from some start positions . For most start states, we collect at least three trajectories starting from there, and hence we can compute a Monte Carlo estimate of the success probability of our policy from those starts . For every  that the estimate is not within the fixed bounds  and , we discard this start so that we don&#8217;t train from it during the next iteration. Starts that were SoID for a previous policy might not be SoID for the current policy because they are now mastered or because the updated policy got worse at them, so it is important to keep filtering out the non-SoID starts to maintain a curriculum adapted to the current agent performance.</p>

<h3>Sampling Nearby</h3>

<p>After filtering the non-SoID we need to obtain new SoIDs to keep expanding the starts from where we train. We do that by sampling states nearby the remaining SoID because those have a similar level of difficulty for the current policy &#8211; and hence might also be SoID. But what is a good way to sample nearby a certain state ? We propose to take random exploratory actions from that , and record the visited states. This technique is preferable to applying noise in state space directly because that might yield states that are not even feasible or that cannot be reached by executing actions from the original .</p>

<h3>Assumptions</h3>

<p>To initialize the algorithm, we need to <strong>seed it with one start at the goal</strong>  and then run brownian motion from it, train from the collected starts, filter out the non-SoID and iterate. This is usually easy to provide when specifying the problem, and is a milder assumption than requiring a full demonstration of how to get to that point.</p>

<p>Our algorithm exploits the capability of <strong>choosing the start distribution</strong> from where the collected trajectories start. This is the case in many systems, like all simulated ones. <a href="http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.7.7601" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Kakade and Langford</a> also build upon this assumption, and propose theoretical evidence of the usefulness of modifying the start distribution.</p>

<h2>Application to Robotics</h2>

<p><em>Navigation</em> to a fixed goal and <em>fine-grained manipulation</em> to a desired configuration are two examples of goal-oriented robotics tasks. We analyze how the proposed algorithm automatically generates a reverse curriculum for the following tasks: Point-mass Maze (Fig. 1a), Ant Maze (Fig. 1b), Ring-on-Peg (Fig. 1c) and Key insertion (Fig. 1d).</p>

<p>
<img height="175" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/point_mass.png" title="point_mass"><img height="175" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ant_maze.png" title="ant_maze"><img height="175" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_on_peg.png" title="ring_on_peg"><img height="175" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/key_insertion.png" title="key_insertion"><br><i>
Fig. 1a-d: Tasks where we illustrate the performance of our method (from left to right): Point-mass Maze, Ant Maze, Ring-on-Peg, Key insertion.
</i>
</p>

<h3>Point-mass Maze</h3>

<p>In this task we want to learn how to reach the end of the red area in the upper-right corner of Fig. 1a from any start point within the maze. We see in Fig. 2 that a randomly initialized policy &#8211; as we have at iteration , has a success probability of zero from everywhere but around the goal. The second row of Fig. 2 shows how our algorithm proposes start positions nearby the goal at . We see in the subsequent columns that the starts generated by our method keep tracking the area where the training policy succeeds sometimes but not always, hence giving a good learning signal for any policy gradient learning method.</p>

<p>
<img width="90%" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/performance.png" title="performance"><br><i>
Fig. 2: Snapshots of the policy performance and the starts generated by our reverse curriculum (replay buffer not depicted for clarity), always tracking the regions at an intermediate level of difficulty.
</i>
</p>

<p>To avoid forgetting how to reach the goal from some areas, we keep a replay buffer of all starts that were SoID for any previous policy. In every training iteration we sample a fraction of the trajectories starting from states in this replay.</p>

<h3>Ant Maze Navigation</h3>

<p>In robotics, a complex coordinated motion is often required to reach the desired configuration. For example, a quadruped like the one in Fig. 1b needs to know how to coordinate all its torques to move and advance towards the goal. As seen in a final policy reported in Video 4, our algorithm is able to learn this behavior even when only a success/failure reward is provided when reaching the goal! The reward function was not modified to include any distance-to-goal, Center of Mass speed, or exploration bonus.</p>

<p>
<img height="250" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ant_maze.gif" title="ant_maze_gif"><br><i>
Video 4: Emergence of complex coordination and use of environment contacts when training with our reverse curriculum method - even under sparse rewards.
</i>
</p>

<h3>Fine-grained Manipulation</h3>

<p>Our method can also tackle complex robotic manipulation problems like the ones depicted in Fig. 1c and 1d. Both task have a seven Degrees of Freedom arm and have complex contact constraints. The first task requires the robot to insert a ring down to the bottom of the peg, and the second task seeks to insert a key in a lock, rotate 90 degrees clockwise, insert it further and rotate 90 degrees counterclockwise. In both cases, a reward is only granted when the desired end configuration is reached. State-of-the-art RL algorithms without curriculum are unable to learn how to solve the task, but with our reverse curriculum generation we can obtain a successful policy from a wide range of start positions, as observed in Videos 5a and 5b.</p>

<p>
<img width="30%" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_success.gif" title="ring_success"><img width="30%" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/key_success.gif" title="key_success"><br><i>
Video 5a (left) and Video 5b (right): Final policies obtained with our reverse curriculum approach in the Ring-on-Peg and the Key Insertion tasks. The agent succeeds from a wide range of initial positions and is able to leverage the contacts to guide itself.
</i>
</p>

<h2>Conclusions and Future Directions</h2>

<p>Recently RL methods have been moving away from the single-task paradigm to tackle sets of tasks. This is an effort to get closer to real-world scenarios, where every time a tasks needs to be executed there are always variations in the starting configuration, goal or other parameters. Therefore it is of utmost importance to advance the field of curriculum learning to exploit the underlying structure of these sets of tasks. Our Reverse Curriculum strategy is a step in this direction, yielding impressive results in locomotion and complex manipulation tasks that cannot be solved without a curriculum.</p>

<p>Furthermore, it can be observed in the videos of our final policy for the manipulation tasks that the agent has learned to exploit the contacts in the environment instead of avoiding them. Therefore, the learning based aspect of the presented method has a great potential to tackle problems that classical motion planning algorithms could struggle with, such as environments with non-rigid objects or with uncertainties in the task geometric parameters. We also leave as future work to combine our curriculum-generation approach with <a href="https://arxiv.org/abs/1703.06907" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">domain randomization methods</a> to obtain policies that are transferable to the real world.</p>

<p>If you want to learn more, check out our paper published in the Conference on Robot Learning:</p>

<p><em>Carlos Florensa, David Held, Markus Wulfmeier, Michael Zhang, Pieter Abbeel. <a href="http://proceedings.mlr.press/v78/florensa17a/florensa17a.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Reverse Curriculum Generation for Reinforcement Learning</a>. In CoRL 2017.</em></p>

<p><em>We have also open-sourced the code in the <a href="https://sites.google.com/view/reversecurriculum" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">project website</a>.</em></p>

<hr><p>We would like to thank the co-authors of this work, who also provided very valuable feedback for this blog post: David Held, Markus Wulfmeier, Michael R. Zhang and Pieter Abbeel.</p>

<h2>Further Readings:</h2>

<p>A. Graves, M. G. Bellemare, J. Menick, R. Munos, and K. Kavukcuoglu. Automated Curriculum Learning for Neural Networks. arXiv preprint, arXiv:1704.03003, 2017.</p>

<p>M. Asada, S. Noda, S. Tawaratsumida, and K. Hosoda. Purposive behavior acquisition for a real robot by Vision-Based reinforcement learning. Machine Learning, 1996.</p>

<p>A. Karpathy and M. Van De Panne. Curriculum learning for motor skills. In Canadian Conference on Artificial Intelligence, pages 325&#8211;330. Springer, 2012.</p>

<p>J. Schmidhuber. POWER PLAY : Training an Increasingly General Problem Solver by Continually Searching for the Simplest Still Unsolvable Problem. Frontiers in Psychology, 2013.</p>

<p>A. Baranes and P.-Y. Oudeyer. Active learning of inverse models with intrinsically motivated goal exploration in robots. Robotics and Autonomous Systems, 61(1), 2013.</p>

<p>S. Sharma and B. Ravindran. Online Multi-Task Learning Using Biased Sampling. arXiv preprint arXiv: 1702.06053, 2017. 9</p>

<p>S. Sukhbaatar, I. Kostrikov, A. Szlam, and R. Fergus. Intrinsic Motivation and Automatic Curricula via Asymmetric Self-Play. arXiv preprint, arXiv: 1703.05407, 2017.</p>

<p>J. A. Bagnell, S. Kakade, A. Y. Ng, and J. Schneider. Policy search by dynamic programming. Advances in Neural Information Processing Systems, 16:79, 2003.</p>

<p>S. Kakade and J. Langford. Approximately Optimal Approximate Reinforcement Learning. International Conference in Machine Learning, 2002.</p>]]></description>
										<content:encoded><![CDATA[<p><strong>By Carlos Florensa</strong></p>
<p>Reinforcement Learning (RL) is a powerful technique capable of solving complex tasks such as <a href="https://arxiv.org/abs/1506.02438" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">locomotion</a>, <a href="https://web.stanford.edu/class/psych209/Readings/MnihEtAlHassibis15NatureControlDeepRL.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Atari games</a>, <a href="https://arxiv.org/abs/1509.02971" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">racing games</a>, and <a href="https://arxiv.org/abs/1504.00702" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">robotic manipulation tasks</a>, all through training an agent to optimize behaviors over a reward function. There are many tasks, however, for which it is hard to design a reward function that is both easy to train and that yields the desired behavior once optimized.<span id="more-94114"></span></p>
<p>Suppose we want a robotic arm to learn how to place a ring onto a peg. The most natural reward function would be for an agent to receive a reward of 1 at the desired end configuration and 0 everywhere else. However, the required motion for this task–to align the ring at the top of the peg and then slide it to the bottom–is impractical to learn under such a binary reward, because the usual random exploration of our initial policy is unlikely to ever reach the goal, as seen in Video 1a. Alternatively, one can try to <a href="https://people.eecs.berkeley.edu/~pabbeel/cs287-fa09/readings/NgHaradaRussell-shaping-ICML1999.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">shape the reward function</a> to potentially alleviate this problem, but finding a good shaping <a href="https://arxiv.org/abs/1704.03073" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">requires considerable expertise and experimentation</a>. For example, directly minimizing the distance between the center of the ring and the bottom of the peg leads to an unsuccessful policy that smashes the ring against the peg, as in Video 1b. We propose a method to learn efficiently without modifying the reward function, by automatically generating a curriculum over start positions.</p>
<table class="col-2">
<tr>
<td style="text-align:center;">
			<img decoding="async" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_fail_cross.gif" alt="ring_fail_cross" />
		</td>
<td style="text-align:center;">
			<img decoding="async" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_shapping_cross.gif" alt="ring_shapping_cross" />
		</td>
</tr>
<tr>
<td>
<p>
      <i>Video 1a: A randomly initialized policy is unable to reach the goal from most start positions, hence being unable to learn.</i>
		</p>
</td>
<td>
<p>
      <i>Video 1b: Shaping the reward with a penalty on the distance from the ring center to the peg bottom yields an undesired behavior.<br />
</i>
		</p>
</td>
</tr>
</table>
<p><!--more--></p>
<h2 id="curriculum-instead-of-reward-shaping">Curriculum instead of Reward Shaping</h2>
<p>We would like to train an agent to reach the goal from any starting position, without requiring an expert to shape the reward. Clearly, not all starting positions are equally difficult. In particular, even a random agent that is placed near to the goal will be able to reach the goal some of the time, receive a reward, and hence start learning! This acquired knowledge can then be bootstrapped to solve the task starting from further away from the goal. By <strong>choosing the ordering of the starting positions that we use in training</strong>, we can exploit this underlying structure of the problem and improve learning efficiency. A key advantage of this technique is that the <strong>reward function is not modified</strong>, and optimizing the sparse reward directly is less prone to yielding undesired behaviors. Ordering a set of related tasks to be learned is referred to as <strong>curriculum learning</strong>, and a central question for us is how to choose this task ordering. Our method, which we explain in more detail below, uses the performance of the learning agent to automatically generate a curriculum of tasks which start from the goal and expand outwards.</p>
<h3 id="reverse-curriculum-intuition">Reverse Curriculum Intuition</h3>
<p>In <em>goal-oriented</em> tasks the aim is to reach a desired configuration from any start state. For example, in the ring-on-peg task introduced above, we desire to place the ring on the peg starting from any configuration. From most start positions, the random exploration of our initial policy never reaches the goal and hence perceives no reward. Nevertheless, it can be seen in Video 2a how a random policy is likely to reach the bottom of the peg if it is initialized from a nearby position. Then, once we have learned how to reach the goal from around the goal, learning from further away is easier since the agent already knows how to proceed if exploratory actions drive its state nearby the goal, as in Video 2b. Eventually, the agent successfully learns to reach  the goal from a wide range of starting positions, as in Video 2c.</p>
<p style="text-align:center;">
<img decoding="async" height="180" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_curr1.gif" title="ring_curr1" /><br />
<img decoding="async" height="180" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_curr2.gif" title="ring_curr2" /><br />
<img decoding="async" height="200" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_curr3.gif" title="ring_curr3" /><br />
<br />
<i><br />
Video 2a-c: Our method strives to learn first starting nearby the goal, and then progressively expands in reverse the positions from where it starts.<br />
</i>
</p>
<p>This method of learning in reverse, or growing outwards from the goal, draws inspiration from Dynamic Programming methods, where the solutions to easier sub-problems are used to compute the solution to harder problems.</p>
<h3 id="starts-of-intermediate-difficulty-soid">Starts of Intermediate Difficulty (SoID)</h3>
<p>To implement this reverse curriculum, we need to ensure that this outwards expansion happens at the right pace for the learning agent. In other words, we want to mathematically describe a set of starts that tracks the current agent performance and provides a good learning signal to our Reinforcement Learning algorithm. In particular we focus on Policy Gradient algorithms, which improve a parameterized policy by taking steps in the direction of an estimated gradient of the total expected reward . This gradient estimation is usually a variation of the original <a href="https://link.springer.com/article/10.1007/BF00992696" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">REINFORCE</a>, which is estimated by collecting $N$ on-policy trajectories <script type="math/tex">\{\tau^i\}_{i=1..N}</script> starting from states <script type="math/tex">\{s^i_0\}_{i=1..N}</script>.</p>
<p><script type="math/tex; mode=display">\nabla_\theta\eta=\frac{1}{N}\sum_{i=1}^N\nabla_\theta\log\pi_\theta(\tau^i)[R(\tau^i, s^i_0)-R(\pi_i, s^i_0)] \quad (1)</script></p>
<p>In <em>goal-oriented</em> tasks, the trajectory reward $R(\tau^i, s^i_0)$ is binary, indicating whether the agent reached the goal. Therefore, the usual baseline $R(\pi_i, s^i_0)$ estimates the probability of reaching the goal if the current policy $\pi_\theta$ is executed starting from $s^i_0$. Hence, we see from Eq. (1) that the terms of the sum corresponding to trajectories collected from starts $s^i_0$ that have success probability 0 or 1 will vanish. These are “wasted” trajectories as they do not contribute to the estimation of the gradient – they are either too hard or too easy. A similar analysis was already introduced in our prior <a href="https://arxiv.org/abs/1705.06366" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">work on multi-task RL</a>. In this case, to avoid training from starts from where our current policy either never gets to the goal or already masters it, we introduce the concept of “Start of Intermediate Difficulty” (SoID), which are start states $s_0$ that satisfy:</p>
<p><script type="math/tex; mode=display">% <![CDATA[
s_0:R_{min} < R(\pi_i, s_0^i) < R_{max} \quad (2) %]]&gt;</script></p>
<p>The values of <script type="math/tex">R_{min}</script> and <script type="math/tex">R_{max}</script> have the straightforward interpretation of minimum success probability acceptable for training from that start and maximum success probability above which we prefer to focus on training from other starts. In all our experiments we used 10% and 90%.</p>
<h2 id="automatic-generation-of-the-reverse-curriculum">Automatic Generation of the Reverse Curriculum</h2>
<p>From the above intuition and derivation, we would like to train our policy with trajectories starting from SoID states. Unfortunately, finding all starts that exactly satisfy Eq. (2) at every policy update is intractable, and hence we introduce an efficient approximation to automatically generate this reverse curriculum: we sample states nearby the starts that were estimated to be SoID during the previous iteration. To do that, we propose a way to <em>filter out non-SoID starts using the trajectories collected during the last training iteration and then sample nearby states</em>. The full algorithm is illustrated in Video 3 and details are given below.</p>
<div class="keep-aspect"><iframe title="Reverse Curriculum Generation for Reinforcement Learning - algorithm sketch" width="500" height="281" src="https://www.youtube-nocookie.com/embed/ANcJ3Hqk7sY?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></div>
<p></p>
<p style="text-align:center;">
<i><br />
Video 3: Animation illustrating the main steps of our algorithm, and how it automatically produces a curriculum adapted to the current agent performance.<br />
</i>
</p>
<h3 id="filtering-out-non-soid">Filtering out Non-SoID</h3>
<p>At every policy gradient training iteration, we collect <script type="math/tex">N</script> trajectories from some start positions <script type="math/tex">\{s^i_0\}_{i=1..N}</script>. For most start states, we collect at least three trajectories starting from there, and hence we can compute a Monte Carlo estimate of the success probability of our policy from those starts <script type="math/tex">R_\theta(s^i_0)</script>. For every <script type="math/tex">s^i_0</script> that the estimate is not within the fixed bounds <script type="math/tex">R_{min}</script> and <script type="math/tex">R_{max}</script>, we discard this start so that we don’t train from it during the next iteration. Starts that were SoID for a previous policy might not be SoID for the current policy because they are now mastered or because the updated policy got worse at them, so it is important to keep filtering out the non-SoID starts to maintain a curriculum adapted to the current agent performance.</p>
<h3 id="sampling-nearby">Sampling Nearby</h3>
<p>After filtering the non-SoID we need to obtain new SoIDs to keep expanding the starts from where we train. We do that by sampling states nearby the remaining SoID because those have a similar level of difficulty for the current policy – and hence might also be SoID. But what is a good way to sample nearby a certain state <script type="math/tex">s^i_0</script>? We propose to take random exploratory actions from that <script type="math/tex">s^i_0</script>, and record the visited states. This technique is preferable to applying noise in state space directly because that might yield states that are not even feasible or that cannot be reached by executing actions from the original <script type="math/tex">s^i_0</script>.</p>
<h3 id="assumptions">Assumptions</h3>
<p>To initialize the algorithm, we need to <strong>seed it with one start at the goal</strong> <script type="math/tex">s^g</script> and then run brownian motion from it, train from the collected starts, filter out the non-SoID and iterate. This is usually easy to provide when specifying the problem, and is a milder assumption than requiring a full demonstration of how to get to that point.</p>
<p>Our algorithm exploits the capability of <strong>choosing the start distribution</strong> from where the collected trajectories start. This is the case in many systems, like all simulated ones. <a href="http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.7.7601" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Kakade and Langford</a> also build upon this assumption, and propose theoretical evidence of the usefulness of modifying the start distribution.</p>
<h2 id="application-to-robotics">Application to Robotics</h2>
<p><em>Navigation</em> to a fixed goal and <em>fine-grained manipulation</em> to a desired configuration are two examples of goal-oriented robotics tasks. We analyze how the proposed algorithm automatically generates a reverse curriculum for the following tasks: Point-mass Maze (Fig. 1a), Ant Maze (Fig. 1b), Ring-on-Peg (Fig. 1c) and Key insertion (Fig. 1d).</p>
<p style="text-align:center;">
<img decoding="async" height="175" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/point_mass.png" title="point_mass" /><br />
<img decoding="async" height="175" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ant_maze.png" title="ant_maze" /><br />
<img decoding="async" height="175" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_on_peg.png" title="ring_on_peg" /><br />
<img decoding="async" height="175" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/key_insertion.png" title="key_insertion" /><br />
<br />
<i><br />
Fig. 1a-d: Tasks where we illustrate the performance of our method (from left to right): Point-mass Maze, Ant Maze, Ring-on-Peg, Key insertion.<br />
</i>
</p>
<h3 id="point-mass-maze">Point-mass Maze</h3>
<p>In this task we want to learn how to reach the end of the red area in the upper-right corner of Fig. 1a from any start point within the maze. We see in Fig. 2 that a randomly initialized policy – as we have at iteration <script type="math/tex">i=1</script>, has a success probability of zero from everywhere but around the goal. The second row of Fig. 2 shows how our algorithm proposes start positions nearby the goal at <script type="math/tex">i=1</script>. We see in the subsequent columns that the starts generated by our method keep tracking the area where the training policy succeeds sometimes but not always, hence giving a good learning signal for any policy gradient learning method.</p>
<p style="text-align:center;">
<img decoding="async" width="90%" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/performance.png" title="performance" /><br />
<br />
<i><br />
Fig. 2: Snapshots of the policy performance and the starts generated by our reverse curriculum (replay buffer not depicted for clarity), always tracking the regions at an intermediate level of difficulty.<br />
</i>
</p>
<p>To avoid forgetting how to reach the goal from some areas, we keep a replay buffer of all starts that were SoID for any previous policy. In every training iteration we sample a fraction of the trajectories starting from states in this replay.</p>
<h3 id="ant-maze-navigation">Ant Maze Navigation</h3>
<p>In robotics, a complex coordinated motion is often required to reach the desired configuration. For example, a quadruped like the one in Fig. 1b needs to know how to coordinate all its torques to move and advance towards the goal. As seen in a final policy reported in Video 4, our algorithm is able to learn this behavior even when only a success/failure reward is provided when reaching the goal! The reward function was not modified to include any distance-to-goal, Center of Mass speed, or exploration bonus.</p>
<p style="text-align:center;">
<img decoding="async" height="250" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ant_maze.gif" title="ant_maze_gif" /><br />
<br />
<i><br />
Video 4: Emergence of complex coordination and use of environment contacts when training with our reverse curriculum method &#8211; even under sparse rewards.<br />
</i>
</p>
<h3 id="fine-grained-manipulation">Fine-grained Manipulation</h3>
<p>Our method can also tackle complex robotic manipulation problems like the ones depicted in Fig. 1c and 1d. Both task have a seven Degrees of Freedom arm and have complex contact constraints. The first task requires the robot to insert a ring down to the bottom of the peg, and the second task seeks to insert a key in a lock, rotate 90 degrees clockwise, insert it further and rotate 90 degrees counterclockwise. In both cases, a reward is only granted when the desired end configuration is reached. State-of-the-art RL algorithms without curriculum are unable to learn how to solve the task, but with our reverse curriculum generation we can obtain a successful policy from a wide range of start positions, as observed in Videos 5a and 5b.</p>
<p style="text-align:center;">
<img decoding="async" width="30%" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/ring_success.gif" title="ring_success" /><br />
<img decoding="async" width="30%" src="http://bair.berkeley.edu/static/blog/reverse_curriculum/key_success.gif" title="key_success" /><br />
<br />
<i><br />
Video 5a (left) and Video 5b (right): Final policies obtained with our reverse curriculum approach in the Ring-on-Peg and the Key Insertion tasks. The agent succeeds from a wide range of initial positions and is able to leverage the contacts to guide itself.<br />
</i>
</p>
<h2 id="conclusions-and-future-directions">Conclusions and Future Directions</h2>
<p>Recently RL methods have been moving away from the single-task paradigm to tackle sets of tasks. This is an effort to get closer to real-world scenarios, where every time a tasks needs to be executed there are always variations in the starting configuration, goal or other parameters. Therefore it is of utmost importance to advance the field of curriculum learning to exploit the underlying structure of these sets of tasks. Our Reverse Curriculum strategy is a step in this direction, yielding impressive results in locomotion and complex manipulation tasks that cannot be solved without a curriculum.</p>
<p>Furthermore, it can be observed in the videos of our final policy for the manipulation tasks that the agent has learned to exploit the contacts in the environment instead of avoiding them. Therefore, the learning based aspect of the presented method has a great potential to tackle problems that classical motion planning algorithms could struggle with, such as environments with non-rigid objects or with uncertainties in the task geometric parameters. We also leave as future work to combine our curriculum-generation approach with <a href="https://arxiv.org/abs/1703.06907" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">domain randomization methods</a> to obtain policies that are transferable to the real world.</p>
<p>If you want to learn more, check out our paper published in the Conference on Robot Learning:</p>
<p><em>Carlos Florensa, David Held, Markus Wulfmeier, Michael Zhang, Pieter Abbeel. <a href="http://proceedings.mlr.press/v78/florensa17a/florensa17a.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Reverse Curriculum Generation for Reinforcement Learning</a>. In CoRL 2017.</em></p>
<p><em>We have also open-sourced the code in the <a href="https://sites.google.com/view/reversecurriculum" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">project website</a>.</em></p>
<hr />
<p>We would like to thank the co-authors of this work, who also provided very valuable feedback for this blog post: David Held, Markus Wulfmeier, Michael R. Zhang and Pieter Abbeel.</p>
<p>This article was initially published on the <a href="“http://bair.berkeley.edu/blog/2017/06/20/welcome/”" data-wpel-link="internal">BAIR blog</a>, and appears here with the authors’ permission.</p>
<h2 id="further-readings">Further Readings:</h2>
<p>A. Graves, M. G. Bellemare, J. Menick, R. Munos, and K. Kavukcuoglu. Automated Curriculum Learning for Neural Networks. arXiv preprint, arXiv:1704.03003, 2017.</p>
<p>M. Asada, S. Noda, S. Tawaratsumida, and K. Hosoda. Purposive behavior acquisition for a real robot by Vision-Based reinforcement learning. Machine Learning, 1996.</p>
<p>A. Karpathy and M. Van De Panne. Curriculum learning for motor skills. In Canadian Conference on Artificial Intelligence, pages 325–330. Springer, 2012.</p>
<p>J. Schmidhuber. POWER PLAY : Training an Increasingly General Problem Solver by Continually Searching for the Simplest Still Unsolvable Problem. Frontiers in Psychology, 2013.</p>
<p>A. Baranes and P.-Y. Oudeyer. Active learning of inverse models with intrinsically motivated goal exploration in robots. Robotics and Autonomous Systems, 61(1), 2013.</p>
<p>S. Sharma and B. Ravindran. Online Multi-Task Learning Using Biased Sampling. arXiv preprint arXiv: 1702.06053, 2017. 9</p>
<p>S. Sukhbaatar, I. Kostrikov, A. Szlam, and R. Fergus. Intrinsic Motivation and Automatic Curricula via Asymmetric Self-Play. arXiv preprint, arXiv: 1703.05407, 2017.</p>
<p>J. A. Bagnell, S. Kakade, A. Y. Ng, and J. Schneider. Policy search by dynamic programming. Advances in Neural Information Processing Systems, 16:79, 2003.</p>
<p>S. Kakade and J. Langford. Approximately Optimal Approximate Reinforcement Learning. International Conference in Machine Learning, 2002.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Towards intelligent industrial co-robots</title>
		<link>https://robohub.org/towards-intelligent-industrial-co-robots/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Mon, 18 Dec 2017 22:34:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://robohub.org/towards-intelligent-industrial-co-robots/</guid>

					<description><![CDATA[Democratization of Robots in Factories

In modern factories, human workers and robots are two major workforces.  For safety concerns, the two are normally separated with robots confined in metal cages, which limits the productivity as well as the flexi...]]></description>
										<content:encoded><![CDATA[<p><strong>By Changliu Liu, Masayoshi Tomizuka </strong></p>
<h2 id="democratization-of-robots-in-factories">Democratization of Robots in Factories</h2>
<p>In modern factories, human workers and robots are two major workforces.  For safety concerns, the two are normally separated with robots confined in metal cages, which limits the productivity as well as the flexibility of production lines. In recent years, attention has been directed to remove the cages so that human workers and robots may collaborate to create a human-robot co-existing factory.<br />
<span id="more-93484"></span></p>
<p>Manufacturers are interested in combining human’s flexibility and robot’s productivity in flexible production lines. The potential benefits of industrial co-robots are huge and extensive, e.g. they may be placed in human-robot teams in flexible production lines, where robot arms and human workers cooperate in handling workpieces, and automated guided vehicles (AGV) co-inhabit with human workers to facilitate factory logistics. In the factories of the future, more and more human-robot interactions are anticipated to take place. Unlike traditional robots that work in structured and deterministic environments, co-robots need to operate in highly unstructured and stochastic environments. The fundamental problem is <em>how to ensure that co-robots operate efficiently and safely in dynamic uncertain environments</em>. In this post, we introduce the robot safe interaction system developed in the <a href="http://msc.berkeley.edu/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Mechanical System Control</a> (MSC) lab.</p>
<p style="text-align:center;">
<img decoding="async" width="50%" src="http://msc.berkeley.edu/assets/images/research/bair/futurefactory.gif" title="future factory" /><img decoding="async" width="50%" src="http://msc.berkeley.edu/assets/images/research/bair/T3.gif" title="future factory" /><br />
<br />
<i><br />
Fig. 1. The factory of the future with human-robot collaborations.<br />
</i>
</p>
<h2 id="existing-solutions">Existing Solutions</h2>
<p>Robot manufacturers including Kuka, Fanuc, Nachi, Yaskawa, Adept and ABB  are providing or working on their solutions to the problem. Several safe cooperative robots or co-robots have been released, such as Collaborative Robots <a href="http://robot.fanucamerica.com/products/robots/collaborative-robot-fanuc-cr-35ia.aspx" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer"><em>CR</em></a> family from FANUC (Japan), <a href="http://www.universalrobots.com/GB/Products.aspx" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">UR5</a> from Universal Robots (Denmark), <a href="http://www.rethinkrobotics.com/products/baxter/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Baxter</a> from Rethink Robotics (US), <a href="http://singularityhub.com/2011/12/09/a-drop-in-solution-for-replacing-humanlabor-kawadas-nextage-robot/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">NextAge</a> from Kawada (Japan) and <a href="http://spectrum.ieee.org/automaton/robotics/industrial-robots/pi4-workerbot-is-one-happy-factory-bot" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">WorkerBot</a> from Pi4_Robotics GmbH (Germany). However, many of these products focus on intrinsic safety, i.e. safety in mechanical design, actuation and low level motion control. Safety during social interactions with humans, which are key to intelligence (including perception, cognition and high level motion planning and control), still needs to be explored.</p>
<h2 id="technical-challenges">Technical Challenges</h2>
<p>Technically, it is challenging to design the behavior of industrial co-robots. In order to make the industrial co-robots human-friendly, they should be equipped with the abilities to: collect environmental data and interpret such data, adapt to different tasks and different environments, and tailor itself to the human workers’ needs. For example, during human-robot collaborative assembly shown in the figure below, the robot should be able to predict that once the human puts the two workpieces together, he will need the tool to fasten the assemble. Then the robot should be able to get the tool and hand it over to the human, while avoid colliding with the human.</p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://msc.berkeley.edu/assets/images/research/bair/int-all.png" title="collaborative assembly" /><br />
<br />
<i><br />
Fig. 2. Human-robot collaborative assembly.<br />
</i>
</p>
<p>To achieve such behavior, the challenges lie in (1) the complication of human behaviors, and (2) the difficulty in assurance of real time safety without sacrificing efficiency. The stochastic nature of human motions brings huge uncertainty to the system, making it hard to ensure safety and efficiency.</p>
<h2 id="the-robot-safe-interaction-system-and-real-time-non-convex-optimization">The Robot Safe Interaction System and Real-time Non-convex Optimization</h2>
<p>The robot safe interaction system (RSIS) has been developed in the <a href="http://msc.berkeley.edu/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Mechanical System Control lab</a>, which establishes a methodology to design the robot behavior to achieve safety and efficiency in peer-to-peer human-robot interactions.</p>
<p>As robots need to interact with humans, who have long acquired interactive behaviors, it is natural to let robot mimic human behavior. Human’s interactive behavior can result from either deliberate thoughts or conditioned reflex. For example, if there is a rear-end collision in the front, the driver of a following car may instinctively hit the brake. However, after a second thought, that driver may speed up to cut into the other lane to avoid chain rear-end. The first is a short-term reactive behavior for safety, while the second needs calculation on current conditions, e.g. whether there is enough space to achieve a full stop, whether there is enough gap for a lane change, and whether it is safer to change lane or do a full stop.</p>
<p>A parallel planning and control architecture has been introduced mimicking these kind of behavior, which included both long term and short term motion planners. The long term planner (efficiency controller) emphasizes efficiency and solves a long-term optimal control problem in receding horizons with low sampling rate. The short term planner (safety controller) addresses real time safety by solving a short-term optimal control problem with high sampling rate based on the trajectories planned by the efficiency controller. This parallel architecture also addresses the uncertainties, where the long term planner plans according to the most-likely behavior of others, and the short term planner considers almost all possible movements of others in the short term to ensure safety.</p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://bair.berkeley.edu/blog/assets/corobots/parallel-architecture.jpg" title="parallel-architecture" /><br />
<br />
<i><br />
Fig. 3. The parallel planning and control architecture in the robot safe interaction system.<br />
</i>
</p>
<p>However, the robot motion planning problems in clustered environment are highly nonlinear and non-convex, hence hard to solve in real time. To ensure timely responses to the change of the environment, fast algorithms are developed for real-time computation, e.g. the convex feasible set algorithm (CFS) for the long term optimization, and the safe set algorithm (SSA) for the short term optimization. These algorithms achieve faster computation by convexification of the original non-convex problem, which is assumed to have convex objective functions, but non-convex constraints. The convex feasible set algorithm (CFS) iteratively solves a sequence of sub-problems constrained in convex subsets of the feasible domain. The sequence of solutions will converge to a local optima. It converges in fewer iterations and run faster than generic non-convex optimization solvers such as sequential quadratic programming (SQP) and interior point method (ITP). On the other hand, the safe set algorithm (SSA) transforms the non convex state space constraints to convex control space constraints using the idea of invariant set.</p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://bair.berkeley.edu/blog/assets/corobots/CFS2.gif" title="CFS" /><br />
<br />
<i><br />
Fig. 4. Illustration of convexification in the CFS algorithm.<br />
</i>
</p>
<p>With the parallel planner and the optimization algorithms, the robot can interact with the environment safely and finish the tasks efficiently.</p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://msc.berkeley.edu/assets/images/research/bair/rsis.gif" title="experiment" /><br />
<br />
<i><br />
Fig. 5. Real time motion planning and control.<br />
</i>
</p>
<h2 id="towards-general-intelligence-the-safe-and-efficient-robot-collaboration-system-serocs">Towards General Intelligence: the Safe and Efficient Robot Collaboration System (SERoCS)</h2>
<p>We now work on an advanced version of RSIS in the Mechanical System Control lab, <a href="http://msc.berkeley.edu/research/serocs.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">the safe and efficient robot collaboration system (SERoCS)</a>, which is supported by National Science Foundation (NSF) <a href="https://www.nsf.gov/awardsearch/showAward?AWD_ID=1734109&amp;HistoricalAwards=false" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Award #1734109</a>. In addition to safe motion planning and control algorithms for safe human-robot interactions (HRI), SERoCS also consists of robust cognition algorithms for environment monitoring, optimal task planning algorithms for safe human-robot collaboration.  The SERoCS will significantly expand the skill sets of the co-robots and prevent or minimize occurrences of human-robot collision and robot-robot collision during operation, hence enables harmonic human-robot collaboration in the future.</p>
<p style="text-align:center;">
<img decoding="async" width="100%" src="http://msc.berkeley.edu/assets/images/research/nri/SERoCS.png" title="Architecture" /><br />
<br />
<i><br />
Fig. 6. SERoCS Architecture.<br />
</i>
</p>
<p>This article was initially published on the <a href="“http://bair.berkeley.edu/blog/2017/06/20/welcome/”" data-wpel-link="internal">BAIR blog</a>, and appears here with the authors’ permission.</p>
<h2 id="references">References</h2>
<table>
<tbody>
<tr>
<td>C. Liu, and M. Tomizuka, “<a href="http://ieeexplore.ieee.org/abstract/document/7487476/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Algorithmic safety measures for intelligent industrial co-robots</a>,” in <em>IEEE International Conference on Robotics and Automation (ICRA)</em>, 2016.</td>
</tr>
<tr>
<td>C. Liu, and M. Tomizuka, “<a href="https://www.springerprofessional.de/en/designing-the-robot-behavior-for-safe-human-robot-interactions/12035766" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Designing the robot behavior for safe human robot interactions</a>”, in <em>Trends in Control and Decision-Making for Human-Robot Collaboration Systems (Y. Wang and F. Zhang (Eds.))</em>. Springer, 2017.</td>
</tr>
<tr>
<td>C. Liu, and M. Tomizuka, “<a href="https://authors.elsevier.com/a/1VlV7c8EXUexT" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Real time trajectory optimization for nonlinear robotic systems: Relaxation and convexification</a>”, in <em>Systems &amp; Control Letters</em>, vol. 108, pp. 56-63, Oct. 2017.</td>
</tr>
<tr>
<td>C. Liu, C. Lin, and M. Tomizuka, “The convex feasible set algorithm for real time optimization in motion planning”,  <a href="https://arxiv.org/abs/1709.00627" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">arXiv:1709.00627</a>.</td>
</tr>
</tbody>
</table>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>FaSTrack: Ensuring safe real-time navigation of dynamic systems</title>
		<link>https://robohub.org/fastrack-ensuring-safe-real-time-navigation-of-dynamic-systems/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Thu, 07 Dec 2017 22:14:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://robohub.org/fastrack-ensuring-safe-real-time-navigation-of-dynamic-systems/</guid>

					<description><![CDATA[
The Problem: Fast and Safe Motion Planning

Real time autonomous motion planning and navigation is hard, especially when we
care about safety.  This becomes even more difficult when we have systems with
complicated dynamics, external disturbances (li...]]></description>
										<content:encoded><![CDATA[<p><iframe width="100%" height="450" src="https://www.youtube-nocookie.com/embed/KcJJOI2TYJA" frameborder="0" allowfullscreen=""></iframe><br />
<strong>By Sylvia Herbert, David Fridovich-Keil, and Claire Tomlin </strong></p>
<h1 id="the-problem-fast-and-safe-motion-planning">The Problem: Fast and Safe Motion Planning</h1>
<p>Real time autonomous motion planning and navigation is hard, especially when we  care about safety.  This becomes even more difficult when we have systems with  complicated dynamics, external disturbances (like wind), and <em>a priori</em> unknown  environments. Our goal in this work is to “robustify” existing real-time motion  planners to guarantee safety during navigation of dynamic systems.</p>
<p>    <span id="more-92897"></span>    </p>
<p>In control theory there are techniques like <a href="http://ieeexplore.ieee.org/abstract/document/1463302/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Hamilton-Jacobi Reachability  Analysis</a> that provide rigorous safety guarantees of system behavior, along  with an optimal controller to reach a given goal (see Fig. 1). However, in  general the computational methods used in HJ Reachability Analysis are only  tractable in decomposable and/or low-dimensional systems; this is due to the  “curse of dimensionality.”  That means for real time planning we can’t process  safe trajectories for systems of more than about two dimensions. Since most  real-world system models like cars, planes, and quadrotors have more than two  dimensions, these methods are usually intractable in real time.</p>
<p>On the other hand, geometric motion planners like rapidly-exploring random trees  (RRT) and model-predictive control (MPC) can plan in real time by using  simplified models of system dynamics and/or a short planning horizon. Although  this allows us to perform real time motion planning, the resulting trajectories  may be overly simplified, lead to unavoidable collisions, and may even be  dynamically infeasible (see Fig. 1).  For example, imagine riding a bike and  following the path on the ground traced by a pedestrian. This path leads you  straight towards a tree and then takes a 90 degree turn away at the last second.  You can’t make such a sharp turn on your bike, and instead you end up crashing  into the tree. Classically, roboticists have mitigated this issue by pretending  obstacles are slightly larger than they really are during planning.  This  greatly improves the chances of not crashing, but still doesn’t provide  guarantees and may lead to unanticipated collisions.</p>
<p>So how do we combine the speed of fast planning with the safety guarantee of  slow planning?</p>
<p style="text-align:center;">  <img decoding="async" src="http://bair.berkeley.edu/blog/assets/fastrack/Figure1.png" width="600" alt="fig1" />  <br />  <i>  Figure 1. On the left we have a high-dimensional vehicle moving through an  obstacle course to a goal. Computing the optimal safe trajectory is a slow and  sometimes intractable task, and replanning is nearly impossible.  On the right  we simplify our model of the vehicle (in this case assuming it can move in  straight lines connected at points).  This allows us to plan very quickly, but  when we execute the planned trajectory we may find that we cannot actually  follow the path exactly, and end up crashing.  </i>  </p>
<h1 id="the-solution-fastrack">The Solution: FaSTrack</h1>
<p>FaSTrack: Fast and Safe Tracking, is a tool that essentially “robustifies” fast  motion planners like RRT or MPC while maintaining real time performance.  FaSTrack allows users to implement a fast motion planner with simplified  dynamics while maintaining safety in the form of a <em>precomputed</em> bound on the  maximum possible distance between the planner’s state and the actual autonomous  system’s state at runtime. We call this distance the <em>tracking error bound</em>.  This precomputation also results in an optimal control lookup table which  provides the optimal error-feedback controller for the autonomous system to  pursue the online planner in real time.</p>
<p style="text-align:center;">  <img decoding="async" src="http://bair.berkeley.edu/blog/assets/fastrack/Figure2.png" height="300" alt="fig2" />  <br />  <i>  Figure 2. The idea behind FaSTrack is to plan using the simplified model (blue),  but precompute a tracking error bound that captures all potential deviations of  the trajectory due to model mismatch and environmental disturbances like wind,  and an error-feedback controller to stay within this bound.  We can then augment  our obstacles by the tracking error bound, which guarantees that our dynamic  system (red) remains safe. Augmenting obstacles is not a new idea in the  robotics community, but by using our tracking error bound we can take into  account system dynamics and disturbances.  </i>  </p>
<h2 id="offline-precomputation">Offline Precomputation</h2>
<p>We precompute this tracking error bound by viewing the problem as a  pursuit-evasion game between a planner and a tracker.  The planner uses a  simplified model of the true autonomous system that is necessary for real time  planning; the tracker uses a more accurate model of the true autonomous system.  We assume that the tracker — the true autonomous system — is always pursuing  the planner. We want to know what the maximum relative distance (i.e. <em>maximum  tracking error</em>) could be in the worst case scenario: when the planner is  actively attempting to evade the tracker.  If we have an upper limit on this  bound then we know the maximum tracking error that can occur at run time.</p>
<p style="text-align:center;">  <img decoding="async" src="http://bair.berkeley.edu/blog/assets/fastrack/Figure3.png" width="500" alt="fig3" />  <br />  <i>  Figure 3. Tracking system with complicated model of true system dynamics  tracking a planning system that plans with a very simple model.  </i>  </p>
<p>Because we care about maximum tracking error, we care about maximum relative  distance.  So to solve this pursuit-evasion game we must first determine the  relative dynamics between the two systems by fixing the planner at the origin  and determining the dynamics of the tracker relative to the planner. We then  specify a cost function as the distance to this origin, i.e. relative distance  of tracker to the planner, as seen in Fig. 4.  The tracker will try to minimize  this cost, and the planner will try to maximize it.  While evolving these  optimal trajectories over time, we capture the highest cost that occurs over the  time period.  If the tracker can always eventually catch up to the planner, this  cost converges to a fixed cost for all time.</p>
<p>The smallest invariant level set of the converged value function provides  determines the tracking error bound, as seen in Fig. 5.  Moreover, the gradient  of the converged value function can be used to create an optimal error-feedback  control policy for the tracker to pursue the planner.  We used <a href="http://www.cs.ubc.ca/~mitchell/ToolboxLS/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Ian Mitchell’s  Level Set Toolbox</a>  and Reachability Analysis to solve this differential  game.  For a more thorough explanation of the optimization, please see <a href="https://arxiv.org/abs/1703.07373" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our  recent paper from the 2017 IEEE Conference on Decision and Control</a>.</p>
<p style="text-align:center;">  <img decoding="async" src="http://bair.berkeley.edu/blog/assets/fastrack/Figure4.gif" height="270" style="margin: 5px;" alt="gif4" />  <img decoding="async" src="http://bair.berkeley.edu/blog/assets/fastrack/Figure5.gif" height="270" style="margin: 5px;" alt="gif5" />  <br />  <i>  Figures 4 &amp; 5: On the left we show the value function initializing at the cost  function (distance to origin) and evolving according to the differential game.  On the right we should 3D and 2D slices of this value function. Each slice can  be thought of as a “candidate tracking error bound.”  Over time, some of these  bounds become infeasible to stay within. The smallest invariant level set of the  converged value function provides us with the tightest tracking error bound that  is feasible.  </i>  </p>
<h2 id="online-real-time-planning">Online real time Planning</h2>
<p>In the online phase, we sense obstacles within a given sensing radius and  imagine expanding these obstacles by the tracking error bound with a Minkowski  sum. Using these padded obstacles, the motion planner decides its next desired  state.  Based on that relative state between the tracker and planner, the  optimal control for the tracker (autonomous system) is determined from the  lookup table.  The autonomous system executes the optimal control, and the  process repeats until the goal has been reached.  This means that the motion  planner can continue to plan quickly, and by simply augmenting obstacles and  using a lookup table for control we can ensure safety!</p>
<p style="text-align:center;">  <img decoding="async" src="http://bair.berkeley.edu/blog/assets/fastrack/Figure6.gif" width="600" alt="gif6" />  <br />  <i>  Figure 6. MATLAB simulation of a 10D near-hover quadrotor model (blue line)  “pursuing” a 3D planning model (green dot) that is using RRT to plan.  As new  obstacles are discovered (turning red), the RRT plans a new path towards the  goal. Based on the relative state between the planner and the autonomous system,  the optimal control can be found via look-up table.  Even when the RRT planner  makes sudden turns, we are guaranteed to stay within the tracking error bound  (blue box).  </i>  </p>
<h1 id="reducing-conservativeness-through-meta-planning">Reducing Conservativeness through Meta-Planning</h1>
<p>One consequence of formulating the safe tracking problem as a pursuit-evasion  game between the planner and the tracker is that the resulting safe tracking  bound is often rather conservative. That is, the tracker can’t <em>guarantee</em> that  it will be close to the planner if the planner is always allowed to do the  <em>worst possible behavior</em>. One solution is to use multiple planning models, each  with its own tracking error bound, simultaneously at planning time. The  resulting “meta-plan” is comprised of trajectory segments computed by each  planner, each labelled with the appropriate optimal controller to track  trajectories generated by that planner. This is illustrated in Fig. 7, where the  large blue error bound corresponds to a planner which is allowed to move very  quickly and the small red bound corresponds to a planner which moves more  slowly.</p>
<p style="text-align:center;">  <img decoding="async" src="http://bair.berkeley.edu/blog/assets/fastrack/Figure7.png" width="500" alt="fig7" />  <br />  <i>  Figure 7. By considering two different planners, each with a different tracking  error bound, our algorithm is able to find a guaranteed safe “meta-plan” that  prefers the less precise but faster-moving blue planner but reverts to the more  precise but slower red planner in the vicinity of obstacles.  This leads to  natural, intuitive behavior that optimally trades off planner conservatism with  vehicle maneuvering speed.  </i>  </p>
<h2 id="safe-switching">Safe Switching</h2>
<p>The key to making this work is to ensure that all transitions between planners  are safe. This can get a little complicated, but the main idea is that a  transition between two planners — call them A and B — is safe if we can  guarantee that the invariant set computed for A is contained within that for B.  For many pairs of planners this is true, e.g. switching from the blue bound to  the red bound in Fig. 7. But often it is not. In general, we need to solve a  dynamic game very similar to the original one in FaSTrack, but where we want to  know the set of states that we will never leave and from which we can guarantee  we end up inside B’s invariant set. Usually, the resulting <em>safe switching  bound</em> (SSB) is slightly larger than A’s tracking error bound (TEB), as shown  below.</p>
<p style="text-align:center;">  <img decoding="async" src="http://bair.berkeley.edu/blog/assets/fastrack/Figure8_v2.png" width="500" alt="fig8" />  <br />  <i>  Figure 8. The safe switching bound for a transition between a planner with a  large tracking error bound to one with a small tracking error bound is generally  larger than the large tracking error bound, as shown.  </i>  </p>
<h2 id="efficient-online-meta-planning">Efficient Online Meta-Planning</h2>
<p>To do this efficiently in real time, we use a modified version of the classical  RRT algorithm. Usually, RRTs work by sampling points in state space and  connecting them with line segments to form a tree rooted at the start point. In  our case, we replace the line segments with the actual trajectories generated by  individual planners. In order to find the shortest route to the goal, we favor  planners that can move more quickly, trying them first and only resorting to  slower-moving planners if the faster ones fail.</p>
<p>We do have to be careful to ensure safe switching bounds are satisfied, however.  This is especially important in cases where the meta-planner decides to  transition to a more precise, slower-moving planner, as in the example above. In  such cases, we implement a one-step virtual backtracking algorithm in which we  make sure the preceding trajectory segment is collision-free using the switching  controller.</p>
<h1 id="implementation">Implementation</h1>
<p>We implemented both FaSTrack and Meta-Planning in C++ / ROS, using low-level  motion planners from the Open Motion Planning Library (OMPL). Simulated results  are shown below, with (right) and without (left) our optimal controller. As you  can see, simply using a linear feedback (LQR) controller (left) provides no  guarantees about staying inside the tracking error bound.</p>
<p style="text-align:center;">  <img decoding="async" src="http://people.eecs.berkeley.edu/~dfk/lqr_video.gif" height="220" style="margin: 5px;" alt="fig09" />  <img decoding="async" src="http://people.eecs.berkeley.edu/~dfk/opt_video.gif" height="220" style="margin: 5px;" alt="fig10" />  <br />  <i>  Figures 9 &amp; 10. (Left) A standard LQR controller is unable to keep the quadrotor  within the tracking error bound. (Right) The optimal tracking controller keeps  the quadrotor within the tracking bound, even during radical changes in the  planned trajectory.  </i>  </p>
<p>It also works on hardware! We tested on the open-source Crazyflie 2.0 quadrotor  platform. As you can see in Fig. 12, we manage to stay inside the tracking bound  at all times, even when switching planners.</p>
<p style="text-align:center;">  <img decoding="async" src="http://bair.berkeley.edu/blog/assets/fastrack/Figure11.png" height="250" style="margin: 5px;" alt="f11" />  <img decoding="async" src="http://bair.berkeley.edu/blog/assets/fastrack/Figure12.png" height="250" style="margin: 5px;" alt="f12" />  <br />  <i>  Figures 11 &amp; 12. (Left) A Crazyflie 2.0 quadrotor being observed by an OptiTrack  motion capture system. (Right) Position traces from a hardware test of the meta  planning algorithm. As shown, the tracking system stays within the tracking  error bound at all times, even during the planner switch that occurs  approximately 4.5 seconds after the start.  </i>  </p>
<p>This article was initially published on the <a href="“http://bair.berkeley.edu/blog/2017/06/20/welcome/”" data-wpel-link="internal">BAIR blog</a>, and appears here with the authors&#8217; permission.</p>
<p>This post is based on the following papers:</p>
<ul>
<li>
<p><strong>FaSTrack: a Modular Framework for Fast and Guaranteed Safe Motion Planning</strong><br />  Sylvia Herbert*, Mo Chen*, SooJean Han, Somil Bansal, Jaime F. Fisac, and Claire J. Tomlin <br />  <a href="https://arxiv.org/abs/1703.07373" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Paper</a>, <a href="http://sylviaherbert.com/fastrack/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Website</a></p>
</li>
<li>
<p><strong>Planning, Fast and Slow: A Framework for Adaptive Real-Time Safe Trajectory Planning</strong><br />  David Fridovich-Keil*, Sylvia Herbert*, Jaime F. Fisac*, Sampada Deglurkar, and Claire J. Tomlin<br />  <a href="https://arxiv.org/abs/1710.04731" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Paper</a>, <a href="https://github.com/HJReachability" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Github</a> (code to appear soon)</p>
</li>
</ul>
<p>We would like to thank our coauthors; developing FaSTrack has been a team effort  and we are incredibly fortunate to have a fantastic set of colleagues on this  project.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>Model-based reinforcement learning with neural network dynamics</title>
		<link>https://robohub.org/model-based-reinforcement-learning-with-neural-network-dynamics/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Fri, 01 Dec 2017 18:09:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://robohub.org/model-based-reinforcement-learning-with-neural-network-dynamics/</guid>

					<description><![CDATA[



Fig 1. A learned neural network dynamics model enables a hexapod robot to learn
to run and follow desired trajectories, using just 17 minutes of real-world
experience.



Enabling robots to act autonomously in the real-world is difficult. Really,
...]]></description>
										<content:encoded><![CDATA[<p><img decoding="async" src="http://robohub.org/wp-content/uploads/2017/11/VelociRoACH.png" alt="" width="1000" height="813" class="aligncenter size-full wp-image-92717" srcset="https://robohub.org/wp-content/uploads/2017/11/VelociRoACH.png 1000w, https://robohub.org/wp-content/uploads/2017/11/VelociRoACH-425x346.png 425w, https://robohub.org/wp-content/uploads/2017/11/VelociRoACH-768x624.png 768w" sizes="(max-width: 1000px) 100vw, 1000px" /><strong>By Anusha Nagabandi and Gregory Kahn</strong></p>
<p>Enabling robots to act autonomously in the real-world is difficult. <a href="https://www.youtube.com/watch?v=g0TaYhjpOfo" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Really, really difficult</a>. Even with expensive robots and teams of world-class researchers, robots still have difficulty autonomously navigating and interacting in complex, unstructured environments.</p>
<p> <span id="more-92662"></span></p>
<p style="text-align:center;"> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_1b.gif" height="240" style="margin: 10px;" alt="fig1b" /> <br /> <i> Fig 1. A learned neural network dynamics model enables a hexapod robot to learn to run and follow desired trajectories, using just 17 minutes of real-world experience. </i> </p>
<p>Why are autonomous robots not out in the world among us? Engineering systems that can cope with all the complexities of our world is hard. From nonlinear dynamics and partial observability to unpredictable terrain and sensor malfunctions, robots are particularly susceptible to Murphy’s law: everything that can go wrong, will go wrong. Instead of fighting Murphy’s law by coding each possible scenario that our robots may encounter, we could instead choose to embrace this possibility for failure, and enable our robots to learn from it. Learning control strategies from experience is advantageous because, unlike hand-engineered controllers, learned controllers can adapt and improve with more data. Therefore, when presented with a scenario in which everything does go wrong, although the robot will still fail, the learned controller will hopefully correct its mistake the next time it is presented with a similar scenario. In order to deal with complexities of tasks in the real world, current learning-based methods often use deep neural networks, which are powerful but not data efficient: These trial-and-error based learners will most often still fail a second time, and a third time, and often thousands to millions of times. The sample inefficiency of modern deep reinforcement learning methods is one of the main bottlenecks to leveraging learning-based methods in the real-world.</p>
<p>We have been investigating sample-efficient learning-based approaches with neural networks for robot control. For complex and contact-rich simulated robots, as well as real-world robots (Fig. 1), our approach is able to learn locomotion skills of trajectory-following using only minutes of data collected from the robot randomly acting in the environment. In this blog post, we’ll provide an overview of our approach and results. More details can be found in our research papers listed at the bottom of this post, including <a href="https://arxiv.org/pdf/1708.02596.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">this paper</a> with <a href="https://github.com/nagaban2/nn_dynamics" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">code here</a>.</p>
<p>  <!--more-->  </p>
<h2 id="sample-efficiency-model-free-versus-model-based">Sample efficiency: model-free versus model-based</h2>
<p>Learning robotic skills from experience typically falls under the umbrella of reinforcement learning. Reinforcement learning algorithms can generally be divided into categories: model-free, which learn a policy or value function, and model-based, which learn a dynamics model. While model-free deep reinforcement learning algorithms are capable of learning a wide range of robotic skills, they typically suffer from <a href="http://www.nature.com/nature/journal/v518/n7540/pdf/nature14236.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">very</a> <a href="https://people.eecs.berkeley.edu/~pabbeel/papers/2015-ICML-TRPO.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">high</a> <a href="https://arxiv.org/pdf/1611.02247.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">sample</a> <a href="https://web.eecs.umich.edu/~baveja/Papers/ICML2016.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">complexity</a>, often requiring millions of samples to achieve good performance, and can typically only learn a single task at a time. Although some prior work has deployed these model-free algorithms for <a href="https://arxiv.org/pdf/1610.00633.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">real-world manipulation tasks</a>, the high sample complexity and inflexibility of these algorithms has hindered them from being widely used to learn locomotion skills in the real world.</p>
<p>Model-based reinforcement learning algorithms are generally regarded as being <a href="http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.436.44&amp;rep=rep1&amp;type=pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">more sample efficient</a>. However, to achieve good sample efficiency, these model-based algorithms have conventionally used either relatively simple <a href="http://papers.nips.cc/paper/5444-learning-neural-network-policies-with-guided-policy-search-under-unknown-dynamics.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">function</a> <a href="http://ieeexplore.ieee.org/document/6907424/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">approximators</a>, which fail to generalize well to complex tasks, or probabilistic dynamics models such as <a href="http://mlg.eng.cam.ac.uk/pub/pdf/DeiRas11.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gaussian</a> <a href="http://ieeexplore.ieee.org/document/7010608/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">processes</a>, which generalize well but have difficulty with complex and high-dimensional domains, such as systems with frictional contacts that induce discontinuous dynamics.  Instead, we use medium-sized neural networks to serve as function approximators that can achieve excellent sample efficiency, while still being expressive enough for generalization and application to various complex and high-dimensional locomotion tasks.</p>
<h2 id="neural-network-dynamics-for-model-based-deep-reinforcement-learning">Neural Network Dynamics for Model-Based Deep Reinforcement Learning</h2>
<p>In our work, we aim to extend the successes that deep neural network models have seen in other domains into model-based reinforcement learning. Prior efforts to combine neural networks with model-based RL in recent years have not achieved the kinds of results that are competitive with simpler models, such as <a href="http://mlg.eng.cam.ac.uk/pub/pdf/DeiRas11.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gaussian processes</a>. For example, <a href="https://arxiv.org/pdf/1603.00748.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Gu et. al.</a> observed that even linear models achieved better performance for synthetic experience generation, while <a href="https://arxiv.org/pdf/1510.09142.pdf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Heess et. al.</a> saw relatively modest gains from including neural network models into a model-free learning system. Our approach relies on a few crucial decisions. First, we use the learned neural network model within a model predictive control framework, in which the system can iteratively replan and correct its mistakes. Second, we use a relatively short horizon look-ahead so that we do not have to rely on the model to make very accurate predictions far into the future. These two relatively simple design decisions enable our method to perform a wide variety of locomotion tasks that have not previously been demonstrated with general-purpose model-based reinforcement learning methods that operate directly on raw state observations.</p>
<p>A diagram of our model-based reinforcement learning approach is shown in Fig. 2. We maintain a dataset of trajectories that we iteratively add to, and we use this dataset to train our dynamics model. The dataset is initialized with random trajectories. We then perform reinforcement learning by alternating between training a neural network dynamics model using the dataset, and using a model predictive controller (MPC) with our learned dynamics model to gather additional trajectories to aggregate onto the dataset. We discuss these two components below.</p>
<p style="text-align:center;"> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_2.png" width="600" alt="fig2" /> <br /> <i> Fig 2. Overview of our model-based reinforcement learning algorithm. </i> </p>
<h3 id="dynamics-model">Dynamics Model</h3>
<p>We parameterize our learned dynamics function as a deep neural network, parameterized by some weights that need to be learned. Our dynamics function takes as input the current state $s_t$ and action $a_t$, and outputs the predicted state difference $s_{t+1}-s_t$. The dynamics model itself can be trained in a supervised learning setting, where collected training data comes in pairs of inputs $(s_t,a_t)$ and corresponding output labels $(s_{t+1},s_t)$.</p>
<p>Note that the “state” that we refer to above can vary with the agent, and it can include elements such as center of mass position, center of mass velocity, joint positions, and other measurable quantities that we choose to include.</p>
<h3 id="controller">Controller</h3>
<p>In order to use the learned dynamics model to accomplish a task, we need to define a reward function that encodes the task. For example, a standard “x_vel” reward could encode a task of moving forward. For the task of trajectory following, we formulate a reward function that incentivizes staying close to the trajectory as well as making forward progress along the trajectory.</p>
<p>Using the learned dynamics model and task reward function, we formulate a model-based controller. At each time step, the agent plans $H$ steps into the future by randomly generating $K$ candidate action sequences, using the learned dynamics model to predict the outcome of those action sequences, and selecting the sequence corresponding to the highest cumulative reward (Fig. 3). We then execute only the first action from the action sequence, and then repeat the planning process at the next time step. This replanning makes the approach robust to inaccuracies in the learned dynamics model.</p>
<p style="text-align:center;"> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_3.png" width="500" alt="fig3" /> <br /> <i> Fig 3. Illustration of the process of simulating multiple candidate action sequences using the learned dynamics model, predicting their outcome, and selecting the best one according to the reward function. </i> </p>
<h2 id="results">Results</h2>
<p>We first evaluated our approach on a variety of MuJoCo agents, including the swimmer, half-cheetah, and ant. Fig. 4 shows that using our learned dynamics model and MPC controller, the agents were able to follow paths defined by a set of sparse waypoints. Furthermore, our approach used only <em>minutes</em> of random data to train the learned dynamics model, showing its sample efficiency.</p>
<p>Note that with this method, we trained the model only once, but simply by changing the reward function, we were able to apply the model at runtime to a variety of different desired trajectories, without a need for separate task-specific training.</p>
<p style="text-align:center;"> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_4a.gif" height="140" style="margin: 6px;" alt="fig4a" /> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_4b.gif" height="140" style="margin: 6px;" alt="fig4b" /> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_4c.gif" height="140" style="margin: 6px;" alt="fig4c" /> <br /> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_4d.gif" height="140" style="margin: 6px;" alt="fig4d" /> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_4e.gif" height="140" style="margin: 6px;" alt="fig4e" /> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_4f.gif" height="140" style="margin: 6px;" alt="fig4f" /> <br /> <i> Fig 4: Trajectory following results with ant, swimmer, and half-cheetah. The dynamics model used by each agent in order to perform these various trajectories was trained just once, using only randomly collected training data. </i> </p>
<p>What aspects of our approach were important to achieve good performance? We first looked at the effect of varying the MPC planning horizon H. Fig. 5 shows that performance suffers if the horizon is too short, possibly due to unrecoverable greedy behavior. For half-cheetah, performance also suffers if the horizon is too long, due to inaccuracies in the learned dynamics model. Fig. 6 illustrates our learned dynamics model for a single 100-step prediction, showing that open-loop predictions for certain state elements eventually diverge from the ground truth. Therefore, an intermediate planning horizon is best to avoid greedy behavior while minimizing the detrimental effects of an inaccurate model.</p>
<p style="text-align:center;"> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_5.png" alt="fig5" /> <br /> <i> Fig 5: Plot of task performance achieved by controllers using different horizon values for planning. Too low of a horizon is not good, and neither is too high of a horizon. </i> </p>
<p style="text-align:center;"> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_6.png" width="600" alt="fig6" /> <br /> <i> Fig 6: A 100-step forward simulation (open-loop) of the dynamics model, showing that open-loop predictions for certain state elements eventually diverge from the ground truth. </i> </p>
<p>We also varied the number of initial random trajectories used to train the dynamics model. Fig. 7 shows that although a higher amount of initial training data leads to higher initial performance, data aggregation allows even low-data initialization experiment runs to reach a high final performance level. This highlights how on-policy data from reinforcement learning can improve sample efficiency.</p>
<p style="text-align:center;"> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_7.png" alt="fig7" /> <br /> <i> Fig 7: Plot of task performance achieved by dynamics models that were trained using differing amounts of initial random data. </i> </p>
<p>It is worth noting that the final performance of the model-based controller is still substantially lower than that of a very good model-free learner (when the model-free learner is trained with thousands of times more experience). This suboptimal performance is sometimes referred to as “model bias,” and is a known issue in model-based RL. To address this issue, we also proposed a hybrid approach that combines model-based and model-free learning to eliminate the asymptotic bias at convergence, though at the cost of additional experience. This hybrid approach, as well as additional analyses, are available in our paper.</p>
<h2 id="learning-to-run-in-the-real-world">Learning to run in the real world</h2>
<p style="text-align:center;"> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_8.png" width="400" alt="fig8" /> <br /> <i> Fig 8: The VelociRoACH is 10 cm in length, approximately 30 grams in weight, can move up to 27 body-lengths per second, and uses two motors to control all six legs. </i> </p>
<p>Since our model-based reinforcement learning algorithm can learn locomotion gaits using orders of magnitude less experience than model-free algorithms, it is possible to evaluate it directly on a real-world robotic platform. In other work, we studied how this method can learn entirely from real-world experience, acquiring locomotion gaits for a millirobot (Fig. 8) completely from scratch.</p>
<p>Millirobots are a promising robotic platform for many applications due to their small size and low manufacturing costs. However, controlling these millirobots is difficult due to their underactuation, power constraints, and size. While hand-engineered controllers can sometimes control these millirobots, they often have difficulties with dynamic maneuvers and complex terrains. We therefore leveraged our model-based learning technique from above to enable the VelociRoACH millirobot to do trajectory following. Fig. 9 shows that our model-based controller can accurately follow trajectories at high speeds, after having been trained using only 17 minutes of random data.</p>
<p style="text-align:center;"> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_9a.gif" height="200" style="margin: 10px;" alt="fig9a" /> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_9b.gif" height="200" style="margin: 10px;" alt="fig9b" /> <br /> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_9c.gif" height="200" style="margin: 10px;" alt="fig9c" /> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/fig_9d.gif" height="200" style="margin: 10px;" alt="fig9d" /> <br /> <i> Fig 9: The VelociRoACH following various desired trajectories, using our model-based learning approach. </i> </p>
<p>To analyze the model’s generalization capabilities, we gathered data on both carpet and styrofoam terrain, and we evaluated our approach as shown in Table 1. As expected, the model-based controller performs best when executed on the same terrain that it was trained on, indicating that the model incorporates knowledge of the terrain. However, performance diminishes when the model is trained on data gathered from both terrains, which likely indicates that more work is needed to develop algorithms for learning models that are effective across various task settings. Promisingly, Table 2 shows that performance increases as more data is used to train the dynamics model, which is an encouraging indication that our approach will continue to improve over time (unlike hand-engineered solutions).</p>
<p style="text-align:center;"> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/table_1.png" width="600" alt="table1" /> <br /> <i> Table 1: Trajectory following costs incurred for models trained with different types of data and for trajectories executed on different surfaces. </i> </p>
<p style="text-align:center;"> <img decoding="async" src="https://people.eecs.berkeley.edu/~nagaban2/misc/bair_blog_figs/table_2.png" width="600" alt="table2" /> <br /> <i> Table 2: Trajectory following costs incurred during the use of dynamics models trained with differing amounts of data. legs. </i> </p>
<p>We hope that these results show the promise of model-based approaches for sample-efficient robot learning and encourage future research in this area.</p>
<p>This article was initially published on the <a href="“http://bair.berkeley.edu/blog/2017/06/20/welcome/”" data-wpel-link="internal">BAIR blog</a>, and appears here with the authors&#8217; permission.</p>
<hr />
<p>We would like to thank Sergey Levine and Ronald Fearing for their feedback.</p>
<p>This post is based on the following papers:</p>
<ul>
<li>
<p><strong>Neural Network Dynamics Models for Control of Under-actuated Legged Millirobots</strong> <br /> A Nagabandi, G Yang, T Asmar, G Kahn, S Levine, R Fearing <br /> <a href="https://arxiv.org/abs/1711.05253" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Paper</a></p>
</li>
<li>
<p><strong>Neural Network Dynamics for Model-Based Deep Reinforcement Learning with Model-Free Fine-Tuning</strong> <br /> A Nagabandi, G Kahn, R Fearing, S Levine <br /> <a href="https://arxiv.org/abs/1708.02596" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Paper</a>, <a href="https://sites.google.com/view/mbmf" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Website</a>, <a href="https://github.com/nagaban2/nn_dynamics" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">Code</a></p>
</li>
</ul>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>DART: Noise injection for robust imitation learning</title>
		<link>https://robohub.org/dart-noise-injection-for-robust-imitation-learning/</link>
		
		<dc:creator><![CDATA[BAIR Blog]]></dc:creator>
		<pubDate>Wed, 22 Nov 2017 00:08:00 +0000</pubDate>
				<category><![CDATA[articles]]></category>
		<guid isPermaLink="false">http://robohub.org/dart-noise-injection-for-robust-imitation-learning/</guid>

					<description><![CDATA[<p>
<img src="http://bair.berkeley.edu/blog/assets/dart/bed_making_gif.gif" alt="Bed-Making GIF" width="600"><br><i>
Toyota HSR Trained with DART to Make a Bed.
</i>
</p>

<p>In Imitation Learning (IL), also known as Learning from Demonstration (LfD), a
robot learns a control policy from analyzing demonstrations of the policy
performed by an algorithmic or human supervisor. For example, to teach a robot
make a bed, a human would tele-operate a robot to perform the task to provide
examples.  The robot then learns a control policy, mapping from images/states to
actions which we hope will generalize to states that were not encountered during
training.</p>

<p>There are two variants of IL: Off-Policy, or Behavior Cloning, where the
demonstrations are given independent of the robot&#8217;s policy.  However, when the
robot encounters novel risky states it may not have learned corrective actions.
This occurs because of &#8220;covariate shift&#8221;  a known challenge, where the states
encountered during training differ from the states encountered during testing,
reducing robustness. Common approaches to reduce covariate shift are On-Policy
methods, such as DAgger, where the evolving robot&#8217;s policy is executed and the
supervisor provides corrective feedback. However, On-Policy methods can be
difficult for human supervisors, potentially dangerous, and computationally
expensive.</p>

<p>This post presents a robust Off-Policy algorithm called DART and summarizes how
injecting noise into the supervisor&#8217;s actions can improve robustness. The
injected noise allows the supervisor to provide corrective examples for the type
of errors the trained robot is likely to make. However, because the optimized
noise is small, it alleviates the difficulties of On-Policy methods. Details on
DART are in a paper that will be presented at <a href="http://www.robot-learning.org/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">the 1st Conference on Robot Learning in
November</a>.</p>

<p>We evaluate DART in  simulation with an algorithmic supervisor on MuJoCo tasks
(Walker, Humanoid, Hopper, Half-Cheetah) and physical experiments with human
supervisors training a Toyota HSR robot to perform grasping in clutter, where a
robot must search through clutter for a goal object.  Finally, we show how
DART can be applied in a complex system that leverages both classical robotics
and learning techniques to teach the first robot to make a bed. For
researchers who want to study and use robust Off-Policy approaches, <strong>we
additionally announce the release of 
<a href="https://berkeleyautomation.github.io/DART/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our codebase</a>
on GitHub</strong>.</p>

<!--more-->

<h1>Imitation Learning&#8217;s Compounding Errors</h1>

<p>In the late 80s, Behavior Cloning was applied to teach cars how to drive, with a
project known as ALVINN (Autonomous Land Vehicle in a Neural Network). In
ALVINN, a neural network was trained on driving demonstrations and learned a
policy that mapped  images of the road to the supervisor&#8217;s steering angle.
Unfortunately, after learning, the policy was unstable, as indicated in the
following video:</p>

<p>
<img src="http://bair.berkeley.edu/blog/assets/dart/alvinn.gif" alt="ALVINN."><br><i>
ALVINN Suffering from Covariate Shift.
</i>
</p>

<p>The car would start drifting to side of the road and not know how to recover.
The reason for the car&#8217;s instability was that no data was collected on the
side of the road.  During the data collection the supervisor always drove along
the center of the road; however, if the robot began to drift from the
demonstrations, it would not know how to recover because it saw no examples.</p>

<p>This example, along with many others that researchers have tried, shows that
Imitation Learning cannot be entirely solved with Behavior Cloning. In
traditional Supervised Learning, the training distribution is de-coupled from
the learned model, whereas in Imitation Learning, <em>the robot&#8217;s policy affects
what state is queried next</em>. Thus the training and testing distributions are no
longer equivalent, and this mismatch is known as 
<strong><a href="http://sifaka.cs.uiuc.edu/jiang4/domain_adaptation/survey/node8.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">covariate shift</a></strong>.</p>

<p>To reduce covariate shift, the objective of Imitation Learning had to be
modified. The robot should now be expected to match the supervisor on the states
it is likely to visit. Thus, if ALVINN is likely to drift to the side of the
road, we expect that it will know what to do in those states.</p>

<p>A robot&#8217;s policy and a supervisor&#8217;s policy can be denoted as $\pi_{\theta}$ and
, where $\pi$ is a function mapping state to action and
$\theta$ is a parametrization, like weights in a neural network.   We can
measure how close two policies are by what actions they apply at a given state,
which we refer to as the surrogate loss, $l$.  A common surrogate loss is the
squared Euclidean distance:</p>

<p>Finally, we need a distribution over trajectories  $p(\xi&#124;\theta)$, which
indicate the trajectories, $\xi$, that are likely under the current policy
$\pi_{\theta}$.  Our objective can then be written as follows:</p>

<p>Hence we want to minimize the expected surrogate loss on the distribution of
states induced by the robot&#8217;s policy. This objective is challenging to solve
because we don&#8217;t know what the robot&#8217;s policy is until after data has been
collected, which creates a <em>chicken and egg</em> situation. We will now discuss an
iterative On-Policy approach to overcome this problem.</p>

<h1>Reducing Shift with On-Policy Methods</h1>

<p>A large body of work from Ross and Bagnell [6,7], has examined the theoretical
consequences of covariate shift. In particular, they proposed the DAgger
algorithm to help correct for it. DAgger can be thought of as an On-Policy
algorithm &#8212; which rolls out the current robot policy during learning.</p>

<p>The key idea of DAgger is to collect data from the current robot policy and
update the model on the aggregate dataset. Implementation of DAgger requires
iteratively rolling out the current robot policy, querying a supervisor for
feedback on the states visited by the robot, and then updating the robot on the
aggregate dataset across all iterations.</p>

<p>
<img src="http://bair.berkeley.edu/blog/assets/dart/DAgger.png" alt="DAGGER." width="600"><br><i>
The DAgger Algorithm.
</i>
</p>

<p>Two years ago, we used DAgger to teach a robot to perform grasping in clutter
(shown below), which requires a robot to search through objects via pushing to
reach a desired goal object. Imitation Learning was advantageous in this task
because we didn&#8217;t need to explicitly model the collision of multiple non-convex
objects.</p>

<p>
<img src="http://bair.berkeley.edu/blog/assets/dart/izzy.gif" alt="Mechnical Search 1." width="250" height="250"><br><i>
Planar Grasping in Clutter.
</i>
</p>

<p>Our planar robot had a neural network policy that mapped images of the workspace
to a control signal. We trained it with DAgger on 160 expert demonstrations.
While we were able to teach the robot how to perform the task with a 90% success
rate, we encountered several major hurdles that made it challenging to increase
the complexity of the task.</p>

<h1>Challenges with On-Policy Methods</h1>

<p>After applying DAgger to teach our robot, we wanted to study and better
understand 3 key limitations related to On-Policy methods in order to scale up
to more challenging tasks.</p>

<h2>Limitation 1: Providing Feedback</h2>

<p>In order to apply feedback to our robot, we had to do so retroactively with a
labeling interface, shown below.</p>

<p>
<img src="http://bair.berkeley.edu/blog/assets/dart/feedback.gif" alt="Mechnical Search 2." width="250" height="250"><br><i>
Supervisor Providing Retroactive Feedback.
</i>
</p>

<p>A supervisor had to manually move the pink overlay to tell the robot what it
should have done after execution. When we tried to retrain the robot with
different supervisors, we found it was very challenging to provide this feedback
for most people. You can think of a human supervisor as a controller that needs
to constantly adjust their actions to obtain the desired effect. However, with
retroactive feedback the human must simulate what the action would be without
seeing the outcome, which is quite unnatural.</p>

<p>To test this hypothesis, we performed a human study with 10 participants to
compare DAgger against Behavior Cloning, where each participant was asked to
train a robot to perform planar part singulation. We found that Behavior Cloning
out-performed DAgger, suggesting that while DAgger mitigates the shift, in
practice it may add systematic noise to the supervisor&#8217;s signal [2].</p>

<h2>Limitation 2: Safety</h2>

<p>On-Policy methods have the additional burden of needing to roll-out the current
robot&#8217;s policy during execution. While our robot was able to perform the task at
the end of training, for most of learning it wasn&#8217;t successful:</p>

<p>
<img src="http://bair.berkeley.edu/blog/assets/dart/policy_gif.gif" alt="Mechnical Search 3." width="250" height="250"><br><i>
Robot Rolling Out Unsuccessful Policy.
</i>
</p>

<p>In unstructured environments, such as a self-driving car or home robotics, this
can be problematic. Ideally, we would like to collect data with the robot while
maintaining high performance throughout the entire process.</p>

<h2>Limitation 3: Computation</h2>

<p>Finally, when building systems either in simulation or the real world, we want
to collect large amounts of data in parallel and update our policy sparingly.
Neural networks can require significant computation time for retraining.
However, On-Policy methods suffer when the policy is not updated frequently
during data collection. Training on a large batch size of new data can cause
significant changes to the current policy, which can push the robot&#8217;s
distribution away from the previously collected data and make the aggregate
dataset stale.</p>

<p>Variants of On-Policy methods have been proposed to solve each of these problems
individually. For example, Ho et al. got rid of the retroactive feedback by
proposing, GAIL, which uses Reinforcement Learning to reduce covariate shift
[8].  Zhang et al. examined how to detect when the policy is about to deviate to
a risky state and asks the supervisor to take over [4]. Sun et al. has explored
incremental gradient updates to the model instead of a full retrain, which is
computationally cheaper [5].</p>

<p>While these methods can each solve some of these problems, ideally we want a
solution to address all three. Off-Policy algorithms like Behavior Cloning do
not exhibit these problems because they passively sample from the supervisor&#8217;s
policy. Thus, we decided instead of extending On-Policy methods it might be more
beneficial to make Off-Policy methods more robust.</p>

<h1>Off-Policy with Noise Injection</h1>

<p>Off-Policy methods, like Behavior Cloning, can in fact have low covariate shift.
If the robot is able to learn the supervisor&#8217;s policy perfectly, then it should
visit the same states as the supervisor. In prior work we empirically found in
simulation that with sufficient data and expressive learners, such as deep
neural networks, Behavior Cloning is at parity with DAgger [2].</p>

<p>In real world domains, though, it is unlikely that a robot can perfectly match a
supervisor. Machine Learning algorithms generally have a long tail in terms of
sample complexity, so the amount of data and computation needed to perfectly
match a supervisor may be unreasonable. However, it is likely that we can
achieve small non-zero test error.</p>

<p>Instead of attempting to perfectly learn the supervisor, we propose simulating
small amounts of error in the supervisor&#8217;s policy to better mimic the trained
robot. Injecting noise into the supervisor&#8217;s policy during teleoperation is one
way to simulate this small test error during data collection. Noise injection
forces the supervisor to provide corrective examples to these small disturbances
as they try to perform the task.  Shown below is the intuition of how noise
injection creates a funnel of corrective examples around the supervisor&#8217;s
distribution.</p>

<p>
<img src="http://bair.berkeley.edu/blog/assets/dart/dart_intuition.png" alt="DART intuition." width="300"><br><i>
Noise Injection forces the supervisor to provide corrective examples,<br>
so that the robot can learn to recover.
</i>
</p>

<p>Additionally, because we are only injecting small noise levels, we don&#8217;t suffer
as many limitations compared to On-Policy methods. A supervisor can normally be
robust to small random disturbances that are concentrated around their current
action. We will now formalize noise injection a bit to help understand its
effect more.</p>

<p>Denote by  a distribution over trajectories with
noise injected into the supervisor&#8217;s distribution
. The parameter $\psi$ represents
the sufficient statistics that define the noise distribution. For example, if
Gaussian noise is injected parameterized by $\psi$, then
.  Note, the stochastic
supervisor&#8217;s distribution is a slight abuse of notation.
 is a distribution over actions,
where as  is a deterministic function mapping to a
single action.</p>

<p>Similar to Behavior Cloning, we can sample demonstrations from the
noise-injected supervisor and minimize the expected loss via standard supervised
learning techniques:</p>

<p>This equation, though, does not explicitly minimize the covariate shift for
arbitrary choices of $\psi$; the $\psi$ needs to be chosen to best simulate the
error of the final robot&#8217;s policy, which may be complex for high dimensional
action spaces.  One approach to choose $\psi$ is grid-search, but this requires
expensive data collection, which can be prohibitive in the physical world or in
high fidelity simulation.</p>

<p>Instead of grid-search, we can formulate the selection of $\psi$ as a maximum
likelihood problem. The objective is to increase the probability of the
supervisor applying the robot&#8217;s control.</p>

<p>This objective states that we want the noise injected supervisor to try and
match the final robot&#8217;s policy. In the paper, we show that this explicitly
minimizes the distance between the supervisor and robot&#8217;s distribution. A clear
limitation of this optimization problem though is that it requires knowing the
final robot&#8217;s distribution $p(\xi&#124;\pi_{\theta^R})$, which is determined only
after the data is collected.  In the next section, we present DART, which
applies an iterative approach to the optimization.</p>

<h2>DART: Disturbances for Augmenting Robot Trajectories</h2>

<p>The above objective cannot be solved because $p(\xi&#124;\pi_{\theta^R})$ is not
known until after the robot has been trained.  We can instead iteratively sample
from the supervisor&#8217;s distribution with the current noise parameter, $\psi_k$,
and minimize the negative log-likelihood of the noise-injected supervisor taking
the current robot&#8217;s, $\pi_{\hat{\theta}}$, control.</p>

<p>The above iterative process can be slow to converge because it is optimizing the
noise with respect to the current robot&#8217;s policy. We can obtain a better
estimate by observing that the supervisor should simulate as much expected error
as the final robot policy, .  It is possible that we have some knowledge of this
quantity from previously training on similar domains. In the paper, we show how
to incorporate this knowledge in the form of a prior. For some common noise
distributions, the objective can be solved in closed form, as detailed in the
paper. Thus, the optimization problem determines the shape of the noise injected
and the prior helps determine the magnitude.</p>

<p>Our algorithm DART, iteratively solves this optimization problem to best set the
noise term. DART is still an iterative algorithm like On-Policy methods.
<em>Through the iterative process, DART optimizes $\psi$ to better simulate the
error in the final robot&#8217;s policy.</em></p>

<h1>Evaluating DART</h1>

<p>To understand how effectively DART reduces covariate shift and to determine if
it suffers from similar limitations as On-Policy methods, we ran experiments in
4 MuJoco domains, as shown below. The supervisor was a policy trained with TRPO
and the noise we injected was Gaussian.</p>

<p>
<img src="http://bair.berkeley.edu/blog/assets/dart/ni_a_results-eps-converted-to.png" alt="DART results."><br></p>

<p>To test if DART suffers from updating the policy after larger batches, we only
updated the model after every $K$ demonstrations for all experiments. DAgger was
updated after every demonstration and DAgger-B was updated after every $K$.  The
results show that DART is able to have the same performance as DAgger, but is
significantly faster in terms of computation.  DAgger-B is relatively similar in
computation time, but suffers significantly in performance, suggesting DART can
significantly reduce computation time.</p>

<p>We finally compared DART to Behavior Cloning in a human study for the task of
grasping in clutter, shown below.  In the task, a Toyota HSR robot was trained to
reach a goal object by pushing objects away with its gripper.</p>

<p>
<img src="http://bair.berkeley.edu/blog/assets/dart/hsr.gif" alt="DART results."><br><i>
Toyota HSR Trained with DART for Grasping in Clutter.
</i>
</p>

<p>The task is more complex than the one above because the robot now sees images of
the world taken from an eye-in-hand camera. We compared 4 humans subjects and
saw that by injecting noise in the controller, we were able to receive a win
over Behavior Cloning of 62%. DART was able to reduce the shift on the task with
human supervisors.</p>

<h1>Robotic Bed Making: A Testbed for Covariate Shift</h1>

<p>To better understand how errors compound in real world robotic systems, we built
a literal test bed. Robotic Bed Making has been a challenging task in robotics
due to it requiring mobile manipulation of deformable objects and sequential
planning. Imitation Learning is one way to sidestep some of the challenges of
deformable object manipulation because it doesn&#8217;t require modeling the bed
sheets.</p>

<p>The goal of our bed making system was to have a robot learn to stretch the
sheets over the bed frame. The task was designed so that the robot must learn
one policy to decide where to grasp the bed sheet and another transition policy
to decide whether the robot should try again or switch to the other bed side.
We trained the bed making policy with 50 demonstrations.</p>

<p>
<img src="http://bair.berkeley.edu/blog/assets/dart/bed_system.png" alt="DART results, bed-making."><br><i>
Bed Making System.
</i>
</p>

<p>DART was applied to inject Gaussian noise into grasping policy because we
assumed there would be considerable error in determining where to grasp. The
optimized covariance matrix decided to inject more noise in the horizontal
direction of the bed, because that is where the edge of the sheet varied more
significantly and subsequently the robot had higher error.</p>

<p>In order to test how large covariate shift was in the system, we can take our
trained policy $\pi_{\theta^R}$ and write its performance with the following
decomposition.</p>

<p>where the first term on the right-hand side corresponds to the covariate shift.
Intuitively, the covariate shift is the difference between the expected error on
the robot&#8217;s distribution and the supervisor&#8217;s distribution. When we measured
these quantities on the bed making setup, we observed noticeable covariate shift
in the transition policy trained with Behavior Cloning.</p>

<p>
<img src="http://bair.berkeley.edu/blog/assets/dart/cs_graph.png" alt="DART results, covariate shift." width="500"><br><i>
Covariate Shift in Bed Making Task.
</i>
</p>

<p>We attribute this covariate shift due to the fact that with Behavior Cloning the
robot rarely saw unsuccessful demonstrations; thus the transition policy never
knew what failure was. DART gave a more diverse set of states, which allowed the
policy to have better class balance. DART was able to train a robust policy that
allowed it to perform the bed making task even when novel objects were placed on
the bed, as shown at the beginning of the blog post. When distractor objects are
placed on the bed DART obtained a 97% sheet coverage, whereas Behavior Cloning
achieved only 63%.</p>

<p>These initial results suggest that covariate shift can occur in modern day
systems that use learning components. We will soon release a longer preprint on
the Bed Making Setup for more information.</p>

<p>DART presents a way to correct for shift via the injection of small optimized
noise. Going forward, we are considering more complex noise models that better
capture the temporal structure of the robot&#8217;s error.</p>

<p>(For papers and updated information, <a href="http://autolab.berkeley.edu/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">see UC Berkeley&#8217;s AUTOLAB website</a>.)</p>

<h2>References</h2>

<ol><li>
    <p>Michael Laskey, Jonathan Lee, Roy Fox, Anca Dragan, Ken Goldberg ; DART:
Noise Injection for Robust Imitation Learning Proceedings of the 1st Annual
Conference on Robot Learning, PMLR 78:143-156, 2017.</p>
  </li>
  <li>
    <p>M. Laskey, C. Chuck, J. Lee, J. Mahler, S. Krishnan, K. Jamieson, A. Dragan,
and K. Goldberg. Comparing human-centric and robot-centric sampling for robot
deep learning from demonstrations. Robotics and Automation (ICRA), 2017 IEEE
International Conference on, pages 358-365. IEEE, 2017</p>
  </li>
  <li>
    <p>M. Laskey, J. Lee, C. Chuck, D. Gealy, W. Hsieh, F. T. Pokorny, A. D. Dragan,
and K. Goldberg. Robot grasping in clutter: Using a hierarchy of supervisors for
learning from demonstrations. In Automation Science and Engineering (CASE), 2016
IEEE International Conference on, pages 827&#8211;834. IEEE, 2016.</p>
  </li>
  <li>
    <p>Zhang, Jiakai, and Kyunghyun Cho. &#8220;Query-Efficient Imitation Learning for
End-to-End Simulated Driving.&#8221; In AAAI, pp. 2891-2897. 2017.</p>
  </li>
  <li>
    <p>W. Sun, A. Venkatraman, G. J. Gordon, B. Boots, and J. A. Bagnell. Deeply
aggrevated: Differentiable imitation learning for sequential prediction.
Proceedings of the 34th International Conference on Machine Learning, PMLR
70:3309-3318, 2017.</p>
  </li>
  <li>
    <p>Ross, St&#233;phane, Geoffrey J. Gordon, and Drew Bagnell. &#8220;A reduction of
imitation learning and structured prediction to no-regret online learning.&#8221;
International Conference on Artificial Intelligence and Statistics. 2011.</p>
  </li>
  <li>
    <p>S. Ross and D. Bagnell. Efficient reductions for imitation learning. In
International Conference on Artificial Intelligence and Statistics, pages
661&#8211;668, 2010.</p>
  </li>
  <li>
    <p>Ho, Jonathan, and Stefano Ermon. &#8220;Generative adversarial imitation learning.&#8221;
Advances in Neural Information Processing Systems. 2016.</p>
  </li>
</ol>]]></description>
										<content:encoded><![CDATA[<div style="width: 706px" class="wp-caption alignnone"><img decoding="async" src="http://bair.berkeley.edu/blog/assets/dart/bed_making_gif.gif" width="696" height="386" class="size-full" /><p class="wp-caption-text">Toyota HSR Trained with DART to Make a Bed.</p></div>
<p><strong>By Michael Laskey, Jonathan Lee, and Ken Goldberg</strong></p>
<p>In Imitation Learning (IL), also known as Learning from Demonstration (LfD), a robot learns a control policy from analyzing demonstrations of the policy performed by an algorithmic or human supervisor. For example, to teach a robot make a bed, a human would tele-operate a robot to perform the task to provide examples.  The robot then learns a control policy, mapping from images/states to actions which we hope will generalize to states that were not encountered during training.</p>
<p>  <span id="more-91776"></span></p>
<p>There are two variants of IL: Off-Policy, or Behavior Cloning, where the demonstrations are given independent of the robot’s policy.  However, when the robot encounters novel risky states it may not have learned corrective actions. This occurs because of “covariate shift”  a known challenge, where the states encountered during training differ from the states encountered during testing, reducing robustness. Common approaches to reduce covariate shift are On-Policy methods, such as DAgger, where the evolving robot’s policy is executed and the supervisor provides corrective feedback. However, On-Policy methods can be difficult for human supervisors, potentially dangerous, and computationally expensive.</p>
<p>This post presents a robust Off-Policy algorithm called DART and summarizes how injecting noise into the supervisor’s actions can improve robustness. The injected noise allows the supervisor to provide corrective examples for the type of errors the trained robot is likely to make. However, because the optimized noise is small, it alleviates the difficulties of On-Policy methods. Details on DART are in a paper that will be presented at <a href="http://www.robot-learning.org/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">the 1st Conference on Robot Learning in November</a>.</p>
<p>We evaluate DART in  simulation with an algorithmic supervisor on MuJoCo tasks (Walker, Humanoid, Hopper, Half-Cheetah) and physical experiments with human supervisors training a Toyota HSR robot to perform grasping in clutter, where a robot must search through clutter for a goal object.  Finally, we show how DART can be applied in a complex system that leverages both classical robotics and learning techniques to teach the first robot to make a bed. For researchers who want to study and use robust Off-Policy approaches, <strong>we additionally announce the release of  <a href="https://berkeleyautomation.github.io/DART/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">our codebase</a> on GitHub</strong>.</p>
<p>  <!--more-->  </p>
<h1 id="imitation-learnings-compounding-errors">Imitation Learning’s Compounding Errors</h1>
<p>In the late 80s, Behavior Cloning was applied to teach cars how to drive, with a project known as ALVINN (Autonomous Land Vehicle in a Neural Network). In ALVINN, a neural network was trained on driving demonstrations and learned a policy that mapped  images of the road to the supervisor’s steering angle. Unfortunately, after learning, the policy was unstable, as indicated in the following video:</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/blog/assets/dart/alvinn.gif" alt="ALVINN." /><br /> <i> ALVINN Suffering from Covariate Shift. </i> </p>
<p>The car would start drifting to side of the road and not know how to recover. The reason for the car’s instability was that no data was collected on the side of the road.  During the data collection the supervisor always drove along the center of the road; however, if the robot began to drift from the demonstrations, it would not know how to recover because it saw no examples.</p>
<p>This example, along with many others that researchers have tried, shows that Imitation Learning cannot be entirely solved with Behavior Cloning. In traditional Supervised Learning, the training distribution is de-coupled from the learned model, whereas in Imitation Learning, <em>the robot’s policy affects what state is queried next</em>. Thus the training and testing distributions are no longer equivalent, and this mismatch is known as  <strong><a href="http://sifaka.cs.uiuc.edu/jiang4/domain_adaptation/survey/node8.html" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">covariate shift</a></strong>.</p>
<p>To reduce covariate shift, the objective of Imitation Learning had to be modified. The robot should now be expected to match the supervisor on the states it is likely to visit. Thus, if ALVINN is likely to drift to the side of the road, we expect that it will know what to do in those states.</p>
<p>A robot’s policy and a supervisor’s policy can be denoted as $\pi_{\theta}$ and <script type="math/tex">\pi_{\theta^*}</script>, where $\pi$ is a function mapping state to action and $\theta$ is a parametrization, like weights in a neural network.   We can measure how close two policies are by what actions they apply at a given state, which we refer to as the surrogate loss, $l$.  A common surrogate loss is the squared Euclidean distance:</p>
<p>  <script type="math/tex; mode=display">l(\pi_{\theta}(x), \pi_{\theta^*}(x)) = \|\pi_{\theta^*}(x) -\pi_{\theta}(x)\|^2_2.</script>  </p>
<p>Finally, we need a distribution over trajectories  $p(\xi|\theta)$, which indicate the trajectories, $\xi$, that are likely under the current policy $\pi_{\theta}$.  Our objective can then be written as follows:</p>
<p>  <script type="math/tex; mode=display">\underset{\theta}{\mbox{min}}\; E_{p(\xi|\theta)} \underbrace{\sum^T_{t=1} l(\pi_{\theta}(x_t), \pi_{\theta^*}(x_t)) }_{J(\theta,\theta^*|\xi)}.</script>  </p>
<p>Hence we want to minimize the expected surrogate loss on the distribution of states induced by the robot’s policy. This objective is challenging to solve because we don’t know what the robot’s policy is until after data has been collected, which creates a <em>chicken and egg</em> situation. We will now discuss an iterative On-Policy approach to overcome this problem.</p>
<h1 id="reducing-shift-with-on-policy-methods">Reducing Shift with On-Policy Methods</h1>
<p>A large body of work from Ross and Bagnell [6,7], has examined the theoretical consequences of covariate shift. In particular, they proposed the DAgger algorithm to help correct for it. DAgger can be thought of as an On-Policy algorithm — which rolls out the current robot policy during learning.</p>
<p>The key idea of DAgger is to collect data from the current robot policy and update the model on the aggregate dataset. Implementation of DAgger requires iteratively rolling out the current robot policy, querying a supervisor for feedback on the states visited by the robot, and then updating the robot on the aggregate dataset across all iterations.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/blog/assets/dart/DAgger.png" alt="DAGGER." width="600" /><br /> <i> The DAgger Algorithm. </i> </p>
<p>Two years ago, we used DAgger to teach a robot to perform grasping in clutter (shown below), which requires a robot to search through objects via pushing to reach a desired goal object. Imitation Learning was advantageous in this task because we didn’t need to explicitly model the collision of multiple non-convex objects.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/blog/assets/dart/izzy.gif" alt="Mechnical Search 1." width="250" height="250" /><br /> <i> Planar Grasping in Clutter. </i> </p>
<p>Our planar robot had a neural network policy that mapped images of the workspace to a control signal. We trained it with DAgger on 160 expert demonstrations. While we were able to teach the robot how to perform the task with a 90% success rate, we encountered several major hurdles that made it challenging to increase the complexity of the task.</p>
<h1 id="challenges-with-on-policy-methods">Challenges with On-Policy Methods</h1>
<p>After applying DAgger to teach our robot, we wanted to study and better understand 3 key limitations related to On-Policy methods in order to scale up to more challenging tasks.</p>
<h2 id="limitation-1-providing-feedback">Limitation 1: Providing Feedback</h2>
<p>In order to apply feedback to our robot, we had to do so retroactively with a labeling interface, shown below.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/blog/assets/dart/feedback.gif" alt="Mechnical Search 2." width="250" height="250" /><br /> <i> Supervisor Providing Retroactive Feedback. </i> </p>
<p>A supervisor had to manually move the pink overlay to tell the robot what it should have done after execution. When we tried to retrain the robot with different supervisors, we found it was very challenging to provide this feedback for most people. You can think of a human supervisor as a controller that needs to constantly adjust their actions to obtain the desired effect. However, with retroactive feedback the human must simulate what the action would be without seeing the outcome, which is quite unnatural.</p>
<p>To test this hypothesis, we performed a human study with 10 participants to compare DAgger against Behavior Cloning, where each participant was asked to train a robot to perform planar part singulation. We found that Behavior Cloning out-performed DAgger, suggesting that while DAgger mitigates the shift, in practice it may add systematic noise to the supervisor’s signal [2].</p>
<h2 id="limitation-2-safety">Limitation 2: Safety</h2>
<p>On-Policy methods have the additional burden of needing to roll-out the current robot’s policy during execution. While our robot was able to perform the task at the end of training, for most of learning it wasn’t successful:</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/blog/assets/dart/policy_gif.gif" alt="Mechnical Search 3." width="250" height="250" /><br /> <i> Robot Rolling Out Unsuccessful Policy. </i> </p>
<p>In unstructured environments, such as a self-driving car or home robotics, this can be problematic. Ideally, we would like to collect data with the robot while maintaining high performance throughout the entire process.</p>
<h2 id="limitation-3-computation">Limitation 3: Computation</h2>
<p>Finally, when building systems either in simulation or the real world, we want to collect large amounts of data in parallel and update our policy sparingly. Neural networks can require significant computation time for retraining. However, On-Policy methods suffer when the policy is not updated frequently during data collection. Training on a large batch size of new data can cause significant changes to the current policy, which can push the robot’s distribution away from the previously collected data and make the aggregate dataset stale.</p>
<p>Variants of On-Policy methods have been proposed to solve each of these problems individually. For example, Ho et al. got rid of the retroactive feedback by proposing, GAIL, which uses Reinforcement Learning to reduce covariate shift [8].  Zhang et al. examined how to detect when the policy is about to deviate to a risky state and asks the supervisor to take over [4]. Sun et al. has explored incremental gradient updates to the model instead of a full retrain, which is computationally cheaper [5].</p>
<p>While these methods can each solve some of these problems, ideally we want a solution to address all three. Off-Policy algorithms like Behavior Cloning do not exhibit these problems because they passively sample from the supervisor’s policy. Thus, we decided instead of extending On-Policy methods it might be more beneficial to make Off-Policy methods more robust.</p>
<h1 id="off-policy-with-noise-injection">Off-Policy with Noise Injection</h1>
<p>Off-Policy methods, like Behavior Cloning, can in fact have low covariate shift. If the robot is able to learn the supervisor’s policy perfectly, then it should visit the same states as the supervisor. In prior work we empirically found in simulation that with sufficient data and expressive learners, such as deep neural networks, Behavior Cloning is at parity with DAgger [2].</p>
<p>In real world domains, though, it is unlikely that a robot can perfectly match a supervisor. Machine Learning algorithms generally have a long tail in terms of sample complexity, so the amount of data and computation needed to perfectly match a supervisor may be unreasonable. However, it is likely that we can achieve small non-zero test error.</p>
<p>Instead of attempting to perfectly learn the supervisor, we propose simulating small amounts of error in the supervisor’s policy to better mimic the trained robot. Injecting noise into the supervisor’s policy during teleoperation is one way to simulate this small test error during data collection. Noise injection forces the supervisor to provide corrective examples to these small disturbances as they try to perform the task.  Shown below is the intuition of how noise injection creates a funnel of corrective examples around the supervisor’s distribution.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/blog/assets/dart/dart_intuition.png" alt="DART intuition." width="300" /> <br /> <i> Noise Injection forces the supervisor to provide corrective examples,<br /> so that the robot can learn to recover. </i> </p>
<p>Additionally, because we are only injecting small noise levels, we don’t suffer as many limitations compared to On-Policy methods. A supervisor can normally be robust to small random disturbances that are concentrated around their current action. We will now formalize noise injection a bit to help understand its effect more.</p>
<p>Denote by <script type="math/tex">p(\xi|\pi_{\theta^*},\psi)</script> a distribution over trajectories with noise injected into the supervisor’s distribution <script type="math/tex">\pi_{\theta^*}(\mathbf{u}|\mathbf{x},\psi)</script>. The parameter $\psi$ represents the sufficient statistics that define the noise distribution. For example, if Gaussian noise is injected parameterized by $\psi$, then <script type="math/tex">\pi_{\theta^*}(\mathbf{u}|\mathbf{x},\psi) = \mathcal{N}(\pi_{\theta^*}(\mathbf{x}), \Sigma)</script>.  Note, the stochastic supervisor’s distribution is a slight abuse of notation. <script type="math/tex">\pi_{\theta^*}(\mathbf{u}|\mathbf{x},\psi)</script> is a distribution over actions, where as <script type="math/tex">\pi_{\theta^*}(\mathbf{x})</script> is a deterministic function mapping to a single action.</p>
<p>Similar to Behavior Cloning, we can sample demonstrations from the noise-injected supervisor and minimize the expected loss via standard supervised learning techniques:</p>
<p>  <script type="math/tex; mode=display">\theta^R = \underset{\theta}{\mbox{argmin }} E_{p(\xi|\pi_{\theta^*},\psi)} J(\theta,\theta^* | \xi)</script>  </p>
<p>This equation, though, does not explicitly minimize the covariate shift for arbitrary choices of $\psi$; the $\psi$ needs to be chosen to best simulate the error of the final robot’s policy, which may be complex for high dimensional action spaces.  One approach to choose $\psi$ is grid-search, but this requires expensive data collection, which can be prohibitive in the physical world or in high fidelity simulation.</p>
<p>Instead of grid-search, we can formulate the selection of $\psi$ as a maximum likelihood problem. The objective is to increase the probability of the supervisor applying the robot’s control.</p>
<p>  <script type="math/tex; mode=display">\underset{\psi}{\mbox{min}} \: E_{p(\xi|\pi_{\theta^R})} -\sum^{T-1}_{t=0} \: \mbox{log} [\pi_{\theta^*}(\pi_{\theta^R}(\mathbf{x_t})|\mathbf{x_t},\psi)]</script>  </p>
<p>This objective states that we want the noise injected supervisor to try and match the final robot’s policy. In the paper, we show that this explicitly minimizes the distance between the supervisor and robot’s distribution. A clear limitation of this optimization problem though is that it requires knowing the final robot’s distribution $p(\xi|\pi_{\theta^R})$, which is determined only after the data is collected.  In the next section, we present DART, which applies an iterative approach to the optimization.</p>
<h2 id="dart-disturbances-for-augmenting-robot-trajectories">DART: Disturbances for Augmenting Robot Trajectories</h2>
<p>The above objective cannot be solved because $p(\xi|\pi_{\theta^R})$ is not known until after the robot has been trained.  We can instead iteratively sample from the supervisor’s distribution with the current noise parameter, $\psi_k$, and minimize the negative log-likelihood of the noise-injected supervisor taking the current robot’s, $\pi_{\hat{\theta}}$, control.</p>
<p>  <script type="math/tex; mode=display">\hat{\psi}_{k+1} = \underset{\psi}{\mbox{argmin}} \: E_{p(\xi|\pi_{\theta^*}, \psi_k)} -\sum^{T-1}_{t=0}\mbox{log} \: [\pi_{\theta^*}(\pi_{\hat{\theta}}(\mathbf{x_t})|\mathbf{x_t},\psi)]</script>  </p>
<p>The above iterative process can be slow to converge because it is optimizing the noise with respect to the current robot’s policy. We can obtain a better estimate by observing that the supervisor should simulate as much expected error as the final robot policy, <script type="math/tex">E_{p(\xi|\pi_{\theta^R})} J(\theta^R,\theta^*|\xi)</script>.  It is possible that we have some knowledge of this quantity from previously training on similar domains. In the paper, we show how to incorporate this knowledge in the form of a prior. For some common noise distributions, the objective can be solved in closed form, as detailed in the paper. Thus, the optimization problem determines the shape of the noise injected and the prior helps determine the magnitude.</p>
<p>Our algorithm DART, iteratively solves this optimization problem to best set the noise term. DART is still an iterative algorithm like On-Policy methods. <em>Through the iterative process, DART optimizes $\psi$ to better simulate the error in the final robot’s policy.</em></p>
<h1 id="evaluating-dart">Evaluating DART</h1>
<p>To understand how effectively DART reduces covariate shift and to determine if it suffers from similar limitations as On-Policy methods, we ran experiments in 4 MuJoco domains, as shown below. The supervisor was a policy trained with TRPO and the noise we injected was Gaussian.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/blog/assets/dart/ni_a_results-eps-converted-to.png" alt="DART results." /> </p>
<p>To test if DART suffers from updating the policy after larger batches, we only updated the model after every $K$ demonstrations for all experiments. DAgger was updated after every demonstration and DAgger-B was updated after every $K$.  The results show that DART is able to have the same performance as DAgger, but is significantly faster in terms of computation.  DAgger-B is relatively similar in computation time, but suffers significantly in performance, suggesting DART can significantly reduce computation time.</p>
<p>We finally compared DART to Behavior Cloning in a human study for the task of grasping in clutter, shown below.  In the task, a Toyota HSR robot was trained to reach a goal object by pushing objects away with its gripper.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/blog/assets/dart/hsr.gif" alt="DART results." /><br /> <i> Toyota HSR Trained with DART for Grasping in Clutter. </i> </p>
<p>The task is more complex than the one above because the robot now sees images of the world taken from an eye-in-hand camera. We compared 4 humans subjects and saw that by injecting noise in the controller, we were able to receive a win over Behavior Cloning of 62%. DART was able to reduce the shift on the task with human supervisors.</p>
<h1 id="robotic-bed-making-a-testbed-for-covariate-shift">Robotic Bed Making: A Testbed for Covariate Shift</h1>
<p>To better understand how errors compound in real world robotic systems, we built a literal test bed. Robotic Bed Making has been a challenging task in robotics due to it requiring mobile manipulation of deformable objects and sequential planning. Imitation Learning is one way to sidestep some of the challenges of deformable object manipulation because it doesn’t require modeling the bed sheets.</p>
<p>The goal of our bed making system was to have a robot learn to stretch the sheets over the bed frame. The task was designed so that the robot must learn one policy to decide where to grasp the bed sheet and another transition policy to decide whether the robot should try again or switch to the other bed side. We trained the bed making policy with 50 demonstrations.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/blog/assets/dart/bed_system.png" alt="DART results, bed-making." /><br /> <i> Bed Making System. </i> </p>
<p>DART was applied to inject Gaussian noise into grasping policy because we assumed there would be considerable error in determining where to grasp. The optimized covariance matrix decided to inject more noise in the horizontal direction of the bed, because that is where the edge of the sheet varied more significantly and subsequently the robot had higher error.</p>
<p>In order to test how large covariate shift was in the system, we can take our trained policy $\pi_{\theta^R}$ and write its performance with the following decomposition.</p>
<p>  <script type="math/tex; mode=display">% <![CDATA[ \begin{align} E_{p(\xi |\pi_{\theta^R})} J(\theta^R,\theta^*|\xi) &#038;= \underbrace{E_{p(\xi |\pi_{\theta^R})}  \sum^T_{t=1} l(\pi_{\theta}(x_t), \pi_{\theta^*}(x_t)) -  E_{p(\xi |\pi_{\theta^*}, \psi)}  \sum^T_{t=1} l(\pi_{\theta}(x_t), \pi_{\theta^*}(x_t))}_{\text{Shift}} \\ &#038;+ \underbrace{E_{p(\xi |\pi_{\theta^*},\psi)}  \sum^T_{t=1} l(\pi_{\theta}(x_t), \pi_{\theta^*}(x_t)) }_{\text{Loss}}, \end{align} %]]&gt;</script>  </p>
<p>where the first term on the right-hand side corresponds to the covariate shift. Intuitively, the covariate shift is the difference between the expected error on the robot’s distribution and the supervisor’s distribution. When we measured these quantities on the bed making setup, we observed noticeable covariate shift in the transition policy trained with Behavior Cloning.</p>
<p style="text-align:center;"> <img decoding="async" src="http://bair.berkeley.edu/blog/assets/dart/cs_graph.png" alt="DART results, covariate shift." width="500" /><br /> <i> Covariate Shift in Bed Making Task. </i> </p>
<p>We attribute this covariate shift due to the fact that with Behavior Cloning the robot rarely saw unsuccessful demonstrations; thus the transition policy never knew what failure was. DART gave a more diverse set of states, which allowed the policy to have better class balance. DART was able to train a robust policy that allowed it to perform the bed making task even when novel objects were placed on the bed, as shown at the beginning of the blog post. When distractor objects are placed on the bed DART obtained a 97% sheet coverage, whereas Behavior Cloning achieved only 63%.</p>
<p>These initial results suggest that covariate shift can occur in modern day systems that use learning components. We will soon release a longer preprint on the Bed Making Setup for more information.</p>
<p>DART presents a way to correct for shift via the injection of small optimized noise. Going forward, we are considering more complex noise models that better capture the temporal structure of the robot’s error.</p>
<p>(For papers and updated information, <a href="http://autolab.berkeley.edu/" data-wpel-link="external" target="_blank" rel="follow external noopener noreferrer">see UC Berkeley’s AUTOLAB website</a>.)</p>
<p>This article was initially published on the <a href="“http://bair.berkeley.edu/blog/2017/06/20/welcome/”" data-wpel-link="internal">BAIR blog</a>, and appears here with the authors&#8217; permission.</p>
<h2 id="references">References</h2>
<ol>
<li>
<p>Michael Laskey, Jonathan Lee, Roy Fox, Anca Dragan, Ken Goldberg ; DART: Noise Injection for Robust Imitation Learning Proceedings of the 1st Annual Conference on Robot Learning, PMLR 78:143-156, 2017.</p>
</li>
<li>
<p>M. Laskey, C. Chuck, J. Lee, J. Mahler, S. Krishnan, K. Jamieson, A. Dragan, and K. Goldberg. Comparing human-centric and robot-centric sampling for robot deep learning from demonstrations. Robotics and Automation (ICRA), 2017 IEEE International Conference on, pages 358-365. IEEE, 2017</p>
</li>
<li>
<p>M. Laskey, J. Lee, C. Chuck, D. Gealy, W. Hsieh, F. T. Pokorny, A. D. Dragan, and K. Goldberg. Robot grasping in clutter: Using a hierarchy of supervisors for learning from demonstrations. In Automation Science and Engineering (CASE), 2016 IEEE International Conference on, pages 827–834. IEEE, 2016.</p>
</li>
<li>
<p>Zhang, Jiakai, and Kyunghyun Cho. “Query-Efficient Imitation Learning for End-to-End Simulated Driving.” In AAAI, pp. 2891-2897. 2017.</p>
</li>
<li>
<p>W. Sun, A. Venkatraman, G. J. Gordon, B. Boots, and J. A. Bagnell. Deeply aggrevated: Differentiable imitation learning for sequential prediction. Proceedings of the 34th International Conference on Machine Learning, PMLR 70:3309-3318, 2017.</p>
</li>
<li>
<p>Ross, Stéphane, Geoffrey J. Gordon, and Drew Bagnell. “A reduction of imitation learning and structured prediction to no-regret online learning.” International Conference on Artificial Intelligence and Statistics. 2011.</p>
</li>
<li>
<p>S. Ross and D. Bagnell. Efficient reductions for imitation learning. In International Conference on Artificial Intelligence and Statistics, pages 661–668, 2010.</p>
</li>
<li>
<p>Ho, Jonathan, and Stefano Ermon. “Generative adversarial imitation learning.” Advances in Neural Information Processing Systems. 2016.</p>
</li>
</ol>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>

<!--
Performance optimized by W3 Total Cache. Learn more: https://www.boldgrid.com/w3-total-cache/?utm_source=w3tc&utm_medium=footer_comment&utm_campaign=free_plugin

Page Caching using Disk: Enhanced 

Served from: robohub.org @ 2026-10-10 23:53:04 by W3 Total Cache
-->