<?xml version="1.0" encoding="UTF-8"?>
<rss  xmlns:atom="http://www.w3.org/2005/Atom" 
      xmlns:media="http://search.yahoo.com/mrss/" 
      xmlns:content="http://purl.org/rss/1.0/modules/content/" 
      xmlns:dc="http://purl.org/dc/elements/1.1/" 
      version="2.0">
<channel>
<title>Computable AI</title>
<link>https://computable.ai/</link>
<atom:link href="https://computable.ai/index.xml" rel="self" type="application/rss+xml"/>
<description>A machine intelligence blog</description>
<generator>quarto-1.7.34</generator>
<lastBuildDate>Sun, 03 Nov 2019 00:00:00 GMT</lastBuildDate>
<item>
  <title>Cox’s Theorem: Establishing Probability Theory</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/coxs-theorem-establishing-probability-theory/</link>
  <description><![CDATA[ 





<section id="ranging-farther-afield" class="level1">
<h1>Ranging farther afield</h1>
<p>Today I’ll be taking advantage of my stated intention to pull back from the stream of <em>recent</em> papers, and look at some papers for their impact or fundamental importance as I see it. So today I’m doing something unusual, highlighting a paper not from last week, but from <em>four years</em> ago, and not directly from AI, but from the field of probability theory: <a href="https://arxiv.org/abs/1507.06597">Cox’s Theorem and the Jaynesian Interpretation of Probability</a>.</p>
<p>I’ve been reading a book by E. T. Jaynes, called <a href="https://www.amazon.com/Probability-Theory-Science-T-Jaynes/dp/0521592712">Probability Theory: The Logic of Science</a>, a brilliant and practical exposition of the Bayesian view of probability theory, partially on <a href="https://www.lesswrong.com/posts/kXSETKZ3X9oidMozA/the-level-above-mine">the recommendation of another AI researcher</a>. The thoughts of an ideal reasoner would have Bayesian structure, so I am both personally and professionally interested in mastering the concepts.</p>
</section>
<section id="overview" class="level1">
<h1>Overview</h1>
<p>Cox’s theorem is an attempt to derive probability theory from a small, common-sense set of uncontroversial desiderata, and to demonstrate its uniqueness as an extension of two-valued (true/false) logic to degrees of belief. That’s a big deal. As today’s paper mentions, Peter Cheeseman <a href="https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1467-8640.1988.tb00091.x">has called</a> Cox’s theorem the “strongest argument for the use of standard (Bayesian) probability theory”. But Cox’s theorem is non-rigorous as originally formulated, and many people have patched up the holes for use in their various fields. Often today, if someone refers to “Cox’s theorem”, they usually mean one of the fixed-up versions.</p>
<p>Jaynes’ version unfortunately contains a mistake, and today’s paper fixes it by replacing some of the axioms with the simple requirement that probability theory remain consistent with respect to repeated events.</p>
<p>It may be difficult without reading the book to see why this paper is important to AI, so perhaps in the near future I’ll discuss that at greater length. For today, however, I’ll simply be explaining each of the axioms, and setting you up to read the paper more easily. It is certainly worth a close reading, to ground your confidence in the interpretation of probability theory as a <em>logical system</em> that extends true-false logic to handle uncertainty, so you can reap the associated benefits.</p>
</section>
<section id="abstract" class="level1">
<h1>Abstract</h1>
<blockquote class="blockquote">
<p>There are multiple proposed interpretations of probability theory: one such interpretation is true-false logic under uncertainty. Cox’s Theorem is a representation theorem that states, under a certain set of axioms describing the meaning of uncertainty, that every true-false logic under uncertainty is isomorphic to conditional probability theory. This result was used by Jaynes to develop a philosophical framework in which statistical inference under uncertainty should be conducted through the use of probability, via Bayes’ Rule. Unfortunately, most existing correct proofs of Cox’s Theorem require restrictive assumptions: for instance, many do not apply even to the simple example of rolling a pair of fair dice. We offer a new axiomatization by replacing various technical conditions with an axiom stating that our theory must be consistent with respect to repeated events. We discuss the implications of our results, both for the philosophy of probability and for the philosophy of statistics.</p>
</blockquote>
<p><img src="https://latex.codecogs.com/png.latex?%5Cnewcommand%7B%5CP%7D%7B%5Cmathbb%7BP%7D%7D%20%5Cnewcommand%7B%5CF%7D%7B%5Cmathscr%7BF%7D%7D"># AxiomsThis paper proposes a new axiomatization of probability theory, with five axioms. As a variant of Cox’s theorem, these axioms are supposed to represent a set of “common sense” desiderata for a logical system under uncertainty. That is, each of these axioms are things we naturally want to be true of any logical system under uncertainty. Cox’s original axioms were more intuitively essential to me, however, so I’ll also try to give justifications for demanding each of the following axioms, as well as explaining them technically.Remember the ultimate goal is to <em>build</em> probability theory up from a minimal set of absolute requirements for <em>any</em> logical system. The punchline is that probability theory as described historically by greats like Kolmogorov turns out to be the <em>unique</em> extension of true-false logic under uncertainty, and we can derive it from “common sense”.To emphasize the point that while we’re writing these axioms we haven’t yet got <em>probability</em>, following Jaynes I’ll refer to our measure of certainty/uncertainty as “plausibility”.</p>
<section id="plausibility-must-be-representable-by-a-real-number" class="level2">
<h2 class="anchored" data-anchor-id="plausibility-must-be-representable-by-a-real-number">1. Plausibility must be representable by a real number</h2>
<blockquote class="blockquote">
<p>Let <img src="https://latex.codecogs.com/png.latex?%5COmega"> be a set and <img src="https://latex.codecogs.com/png.latex?%5Cmathscr%7BF%7D"> be a <img src="https://latex.codecogs.com/png.latex?%5Csigma">-algebra on <img src="https://latex.codecogs.com/png.latex?%5COmega">.</p>
<p>Let <img src="https://latex.codecogs.com/png.latex?%5CP:%20%5CF%20%5Ctimes%20(%5CF%20%5Csetminus%20%5Cemptyset)%20%5Crightarrow%20R%20%5Csubseteq%20%5Cmathbb%7BR%7D"> be a function, written using notation <img src="https://latex.codecogs.com/png.latex?%5CP(A%7CB)">.</p>
</blockquote>
<p>It makes intuitive sense that we should be able to measure our uncertainty on a smooth, finite scale, so it makes sense to demand that our plausibility scale be chosen from some definite subset of the reals.</p>
<p><img src="https://latex.codecogs.com/png.latex?%5CF"> being “<a href="https://en.wikipedia.org/wiki/Sigma-algebra">a <img src="https://latex.codecogs.com/png.latex?%5Csigma">-algebra on <img src="https://latex.codecogs.com/png.latex?%5COmega"></a>” means that it is the set of every subset of <img src="https://latex.codecogs.com/png.latex?%5COmega"> (including <img src="https://latex.codecogs.com/png.latex?%5COmega"> and <img src="https://latex.codecogs.com/png.latex?%5Cemptyset">), is closed under complement, and is closed under countable unions. (Being “closed under” some operation means that taking that operation on any element in the set yields an element that’s also defined to be in the set.) The idea is that <img src="https://latex.codecogs.com/png.latex?%5COmega"> comprises all primitive events, and <img src="https://latex.codecogs.com/png.latex?%5CF"> therefore includes every possible logical combination of these primitive events, in a way that makes it eqivalent to a Boolean algebra.</p>
<p>I found it clarifying that <img src="https://latex.codecogs.com/png.latex?%5CP(%5COmega)=1">. That’s what made it click for me that a set in <img src="https://latex.codecogs.com/png.latex?%5CF"> represents a disjunction of primitive events, and <img src="https://latex.codecogs.com/png.latex?%5COmega"> contains <em>all</em> primitive events, so <img src="https://latex.codecogs.com/png.latex?%5CP(%5COmega)"> is the probability that <em>anything</em> happens.</p>
<p><img src="https://latex.codecogs.com/png.latex?%5CP(A%7CB)"> is a function of two arguments <img src="https://latex.codecogs.com/png.latex?A,B%20%5Cin%20%5Cmathscr%7BF%7D">, and B cannot be empty. The interpretation is, “The probability of some event A, given that event B is true.” The second argument cannot be empty, Jaynes often describes it as “the background information”, including everything else known (such as the rules of probability themselves, and the number of penguins in Antarctica).</p>
<p>The arguments of <img src="https://latex.codecogs.com/png.latex?%5CP"> are sets, but as the paper mentions, “by <a href="https://www.jstor.org/stable/1989664">Stone’s Representation Theorem</a>, every Boolean algebra is isomorphic to an algebra of sets”.</p>
</section>
<section id="sequential-continuity" class="level2">
<h2 class="anchored" data-anchor-id="sequential-continuity">2. Sequential continuity</h2>
<blockquote class="blockquote">
<p>We have that <img src="https://latex.codecogs.com/png.latex?A_1%20%5Csubseteq%20A_2%20%5Csubseteq%20A_3%20%5Csubseteq%5Cldots%20%5Ctext%7B%20such%20that%20%7D%20A_i%20%5Cnearrow%20A%20%5Ctext%7B%20implies%20%7D%20%5CP%20(A_i%20%7C%20B)%5Cnearrow%20%5CP(A%20%7C%20B%20)"> for all <img src="https://latex.codecogs.com/png.latex?A,%20A_i,%20B">.</p>
</blockquote>
<p>Another intuitive requirement for a system of logical inference is that our plausibility measure return arbitrarily small differences in plausibility for arbitrarily small changes in truth value. This concept is also known as “continuity”.</p>
<p>If you can arrange a sequence of events (sets) so that earlier events (e.g., <img src="https://latex.codecogs.com/png.latex?A_1">) are included in later events (e.g., <img src="https://latex.codecogs.com/png.latex?A_3">), then there is “sequential continuity” between earlier sets and later sets in this sequence. In the notation of the paper, <img src="https://latex.codecogs.com/png.latex?A_1%20%5Cnearrow%20A_3">.</p>
<p>What this axiom is saying is that as long as there is sequential continuity between two logical propositions, there is also sequential continuity between their plausibilities. This formalizes our requirement for continuity. Also notice that if <img src="https://latex.codecogs.com/png.latex?%5CP%20(A_i%20%7C%20B)%5Cnearrow%20%5Cmathbb%7BP%7D(A%20%7C%20B%20)"> then <img src="https://latex.codecogs.com/png.latex?%5CP%20(A_i%20%7C%20B)%20%5Cleq%20%5Cmathbb%7BP%7D(A%20%7C%20B%20)">, because our definition of sequential continuity also implies that the cardinality of the sets is non-decreasing. This will be useful reading the proof.</p>
</section>
<section id="decomposability" class="level2">
<h2 class="anchored" data-anchor-id="decomposability">3. Decomposability</h2>
<blockquote class="blockquote">
<p><img src="https://latex.codecogs.com/png.latex?%5CP(AB%20%7C%20C%20)"> can be written as <img src="https://latex.codecogs.com/png.latex?%5CP(A%20%7C%20C%20)%20%5Ccirc%20%5CP(B%20%7C%20AC)"> for some some function <img src="https://latex.codecogs.com/png.latex?%5Ccirc%20:%20(R%20%5Ctimes%20R)%20%5Crightarrow%20R">.</p>
</blockquote>
<p>This is the first axiom that I had trouble seeing as intuitive, and in fact I thought it was a bit question-begging at first because it looks like the product rule. It represents the demand that plausibilities of compound propositions be decomposable into plausibilities of the their constituents, and that that decomposition has a particular form. It’s the demand that it follow a particular form that seems somewhat arbitrary to me at first. Of course we would want to be able to decompose compound uncertainty into more fundamental elements, or else probability theory wouldn’t be very useful. But why should it take the form described of <img src="https://latex.codecogs.com/png.latex?%5Ccirc">?</p>
<p>The answer is that this form is <em>minimal</em> for decomposability. That is, it’s the weakest statement that could be made about the details of decomposition. In English: “The plausibility of A <em>and</em> B is a function of the plausibility of one of those (say, <img src="https://latex.codecogs.com/png.latex?A">), and the plausibility of the other (<img src="https://latex.codecogs.com/png.latex?B">) once we can assume <img src="https://latex.codecogs.com/png.latex?A"> is true.”</p>
<p>Note that logical conjunctions are commutative (<img src="https://latex.codecogs.com/png.latex?AB%20=%20BA">), so by this axiom <img src="https://latex.codecogs.com/png.latex?%5CP(AB%20%7C%20C%20)"> can <em>also</em> be written as <img src="https://latex.codecogs.com/png.latex?%5CP(B%20%7C%20C%20)%20%5Ccirc%20%5CP(A%20%7C%20BC)">. They prove later also that <img src="https://latex.codecogs.com/png.latex?%5Ccirc"> is commutative, but that is not assumed in the axioms.</p>
</section>
<section id="negation" class="level2">
<h2 class="anchored" data-anchor-id="negation">4. Negation</h2>
<blockquote class="blockquote">
<p>There exists a function <img src="https://latex.codecogs.com/png.latex?N%20:%20R%20%5Crightarrow%20R"> such that <img src="https://latex.codecogs.com/png.latex?%0A%5CP(A%5Ec%20%7C%20B)=%20N%5B%20%5CP(A%20%7C%20B)%5D%0A"> for all <img src="https://latex.codecogs.com/png.latex?A,B">.</p>
</blockquote>
<p>This axiom also seemed a bit question-begging to me, because it looks like the sum rule of probability theory, and because it seemed arbitrary that you would want uniquely determined probabilities for the negations of propositions.</p>
<p>Upon further reflection, however, this seems like a reasonable demand to be consistent with two-valued logic. Every proposition <img src="https://latex.codecogs.com/png.latex?A"> in true-false logic has a unique proposition <img src="https://latex.codecogs.com/png.latex?A%5Ec"> representating its negation, (This superscript complement notation emphasizes the representation as propositions as sets, but is equivalent to <img src="https://latex.codecogs.com/png.latex?%5Cbar%20A">, <img src="https://latex.codecogs.com/png.latex?%5Cneg%20A">, etc.) so it makes sense that an extension of true-false logic to uncertainty would also include a method of determining the opposite.</p>
<p>In actual fact, this <em>may</em> be the most controversial axiom, since there are logics other than true-false logic that don’t require the “law of the excluded middle” (they allow “maybe”). But if you are willing to accept that all well-formed propositions are either true or false, and our system of plausibility represents levels of certainty about their truth or falsehood, then this axiom represents a reasonable and necessary demand.</p>
</section>
<section id="consistency-under-extension" class="level2">
<h2 class="anchored" data-anchor-id="consistency-under-extension">5. Consistency under extension</h2>
<blockquote class="blockquote">
<p>If <img src="https://latex.codecogs.com/png.latex?(%5COmega,%20%5Cmathscr%7BF%7D,%20%5CP)"> satisfies the axioms above, then <img src="https://latex.codecogs.com/png.latex?(%5COmega%20%5Ctimes%20%5COmega,%20%5Cmathscr%7BF%7D%20%5Cotimes%20%5Cmathscr%7BF%7D,%20%5CP%20%5Coperatorname%7B%5Ccirc%7D%20%5CP)"> must as well, i.e., the definition <img src="https://latex.codecogs.com/png.latex?%5CP(A%20%5Ctimes%20B%20%7C%20C%20%5Ctimes%20D)%20=%20%5CP(A%20%7C%20C)%20%5Ccirc%20%5CP(B%20%7C%20D)"> is consistent.</p>
</blockquote>
<p>This axiom represents the core of the authors’ contribution. Although there were many correct variants of Cox’s theorem, and many ways to axiomatize probability theory, they all had either disappointingly narrow scope, or had lost their intuitive nature in the formalization. The authors’ of our paper replace several technical axioms from other axiomatizations with this one demand <em>that their rules be consistent under extention to repeated events</em>.</p>
<p>In English, this axiom is, “If the rules apply to a single trial (e.g., a single coinflip), then they also apply to a system of two independent trials (e.g., two coinflips).” To me, that’s obviously intuitive, so it’s delightful to find that it covers so much ground.</p>
<p>Examining their formal expression, with the coinflips example, with <img src="https://latex.codecogs.com/png.latex?A"> meaning “heads on the first coinflip” and B meaning “tails on the second coinflip”:</p>
<p><img src="https://latex.codecogs.com/png.latex?%5CP(A%20%5Ctimes%20B%20%7C%20C%20%5Ctimes%20D)"> means “the plausibility of heads-then-tails given two piles of background information <img src="https://latex.codecogs.com/png.latex?C"> and <img src="https://latex.codecogs.com/png.latex?D">”. The axiom states this must equal <img src="https://latex.codecogs.com/png.latex?%5CP(A%20%7C%20C)%20%5Ccirc%20%5CP(B%20%7C%20D)">, meaning that the plausibility of a pair of coinflips coming up heads-tails is equal to the plausibility of a single coinflip coming up heads (given background information <img src="https://latex.codecogs.com/png.latex?C">), composed (using <img src="https://latex.codecogs.com/png.latex?%5Ccirc">) with another coinflip coming up tails (given background information <img src="https://latex.codecogs.com/png.latex?D">).</p>
</section>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li>I hope this exposition of the axioms helps you read the paper yourself, though I realize I may not have provided sufficient motivation to do so yet. That would make it a bit like <a href="https://computable.ai/articles/2019/Mar/10/boltzmann-machines-differentiation-work.html">my post deriving something surprising about Boltzmann machines</a> without first explaining what Boltzmann machines <em>are</em>. I intend to rectify this in the future for both posts.</li>
<li>I could make this a lot clearer for people with less set theory, group theory, or probability theory background. If that would be helpful to you, please leave me a comment on what specifically didn’t make sense so I can get a feel for my audience.</li>
<li>To memorize these and make reading the proof easier, I labeled each of the five axioms with some relevant symbol, and combined them into a mneumonic. In case that helps you too, here it is: <img src="https://latex.codecogs.com/png.latex?%5Cmathbb%7BR%7D"> <img src="https://latex.codecogs.com/png.latex?%5Cnearrow"> <img src="https://latex.codecogs.com/png.latex?%5Ccirc"> <img src="https://latex.codecogs.com/png.latex?N"> <img src="https://latex.codecogs.com/png.latex?%5Ctimes">.</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/coxs-theorem-establishing-probability-theory/</guid>
  <pubDate>Sun, 03 Nov 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/coxs-theorem-establishing-probability-theory.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>Comments on Eight Abstracts</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/comments-on-eight-abstracts/</link>
  <description><![CDATA[ 





<section id="this-last-week" class="level1">
<h1>This (last) week</h1>
<p>Alas, I bit off more than I could chew last week. You’ll see what I mean in a moment. However, I’ve decided to define the problem away, as part of an effort to more effectively juggle all of my life responsibilities:</p>
<p><strong>ArXiv Highlights will be bi-weekly</strong> from here on out. I’m also going to be a <em>little</em> less strict about when I sample papers from, so that I don’t feel so constrained to do “last week’s” arXiv announcements. The attentive reader may have noticed that I’ve already occasionally sampled from outside of the week’s announcements, and I’d actually prefer to do that more often so that I can hit <em>key</em> papers instead of just <em>new</em> papers.</p>
<hr>
<p>I couldn’t just pick one paper last week, since so many seemed relevant and interesting. Therefore I’m experimenting with yet another format for arXiv highlights: posting all the abstracts, and commenting a bit on each one. The goal is to work each of these concepts into my memory (and yours) so that they’ll spring to mind when we need them.</p>
<p>In arXiv announcement order:</p>
<ol type="1">
<li><a href="https://arxiv.org/abs/1909.07528v1">Emergent Tool Use From Multi-Agent Autocurricula</a></li>
<li><a href="https://arxiv.org/abs/1909.10618v1">Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?</a></li>
<li><a href="https://arxiv.org/abs/1909.11145v1">Brain-Inspired Hardware for Artificial Intelligence: Accelerated Learning in a Physical-Model Spiking Neural Network</a></li>
<li><a href="https://arxiv.org/abs/1909.11373v1">Pre-training as Batch Meta Reinforcement Learning with tiMe</a></li>
<li><a href="https://arxiv.org/abs/1907.06511v2">Reinforcement Learning with Chromatic Networks</a></li>
<li><a href="https://arxiv.org/abs/1907.08225v2">Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery</a></li>
<li><a href="https://arxiv.org/abs/1909.11821v1">Model Imitation for Model-Based Reinforcement Learning</a></li>
<li><a href="https://arxiv.org/abs/1907.08591v2">Zermelo’s problem: Optimal point-to-point navigation in 2D turbulent flows using Reinforcement Learning</a></li>
</ol>
</section>
<section id="emergent-tool-use-from-multi-agent-autocurricula" class="level1">
<h1>1. Emergent Tool Use From Multi-Agent Autocurricula</h1>
<blockquote class="blockquote">
<p>Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a self-supervised autocurriculum inducing multiple distinct rounds of emergent strategy, many of which require sophisticated tool use and coordination. We find clear evidence of six emergent phases in agent strategy in our environment, each of which creates a new pressure for the opposing team to adapt; for instance, agents learn to build multi-object shelters using moveable boxes which in turn leads to agents discovering that they can overcome obstacles using ramps. We further provide evidence that multi-agent competition may scale better with increasing environment complexity and leads to behavior that centers around far more human-relevant skills than other self-supervised reinforcement learning methods such as intrinsic motivation. Finally, we propose transfer and fine-tuning as a way to quantitatively evaluate targeted capabilities, and we compare hide-and-seek agents to both intrinsic motivation and random initialization baselines in a suite of domain-specific intelligence tests.</p>
</blockquote>
<p>https://arxiv.org/abs/1909.07528v1</p>
<p>I notice that OpenAI, DeepMind, and Google Brain are involved in a lot of the interesting work in reinforcement learning lately, and many of the papers that catch my eye have at least some authors from either of these organizations. I’m an aspiring Bayesian, so it wasn’t <em>too</em> long before I starting reading papers <em>because</em> they were authored by one of these organizations.</p>
<p>Anyway, the term “autocurriculum” seems to come from <a href="https://arxiv.org/abs/1903.00742">this DeepMind paper</a>:</p>
<blockquote class="blockquote">
<p>Here we explore the hypothesis that multi-agent systems sometimes display intrinsic dynamics arising from competition and cooperation that provide a naturally emergent curriculum, which we term an autocurriculum.</p>
</blockquote>
<p>This gives me a word for something I’ve observed about my young son: The activities he’s naturally inclined to engage in at each stage of his development seem uncannily well-suited for teaching him the <em>next</em> thing he should learn. Wanting to put things in his mouth, plus a capacity for boredom, motivated him to develop reaching and grabbing, then crawling, then pathfinding, then complex navigation…</p>
<p>In the case of the OpenAI paper, putting multiple adversarial agents into complex environments and allowing them to learn causes them to learn new behaviors <em>in phases</em>.</p>
<blockquote class="blockquote">
<p>We find clear evidence of six emergent phases in agent strategy in our environment, each of which creates a new pressure for the opposing team to adapt</p>
</blockquote>
<p>Each time a team of agents learns a dominant strategy, the opposing team is pressured to develop a strategy capable of defeating it, which then pressures the first team to come up with <em>another</em> stretegy to defeat <em>that</em> one, and on and on until a truly dominant strategy emerges.</p>
<p>My takeaway from this is that emergent autocurricula may be another good reason for me to study multi-agent systems.</p>
<p>This paper comes with a nice blog post and a cute video: https://openai.com/blog/emergent-tool-use/</p>
</section>
<section id="why-does-hierarchy-sometimes-work-so-well-in-reinforcement-learning" class="level1">
<h1>2. Why Does Hierarchy (Sometimes) Work So Well in Reinforcement Learning?</h1>
<blockquote class="blockquote">
<p>Hierarchical reinforcement learning has demonstrated significant success at solving difficult reinforcement learning (RL) tasks. Previous works have motivated the use of hierarchy by appealing to a number of intuitive benefits, including learning over temporally extended transitions, exploring over temporally extended periods, and training and exploring in a more semantically meaningful action space, among others. However, in fully observed, Markovian settings, it is not immediately clear why hierarchical RL should provide benefits over standard “shallow” RL architectures. In this work, we isolate and evaluate the claimed benefits of hierarchical RL on a suite of tasks encompassing locomotion, navigation, and manipulation. Surprisingly, we find that most of the observed benefits of hierarchy can be attributed to improved exploration, as opposed to easier policy learning or imposed hierarchical structures. Given this insight, we present exploration techniques inspired by hierarchy that achieve performance competitive with hierarchical RL while at the same time being much simpler to use and implement.</p>
</blockquote>
<p>https://arxiv.org/abs/1909.10618v1</p>
<p>The big finding here is that “most of the observed benefits of hierarchy can be attributed to improved exploration”.</p>
<p>This is not the first time I’ve heard that a complicated technique in RL has been studied and found to boil down to better exploration or more even coverage of the state space. <a href="https://arxiv.org/abs/1902.10250">Diagnosing Bottlenecks in Deep Q-learning Algorithms</a> contained a similar revelation about replay buffer sampling, for example, and I get the impression that the maximum entropy RL framework seems to be overtaking more ad hoc trust region methods such as PPO. Anyway, this is why theory is important even to mere industry practitioners. As theory catches up to practice, we learn <em>why</em> things work, and the answers are often surprising and useful.</p>
</section>
<section id="brain-inspired-hardware-for-artificial-intelligence-accelerated-learning-in-a-physical-model-spiking-neural-network" class="level1">
<h1>3. Brain-Inspired Hardware for Artificial Intelligence: Accelerated Learning in a Physical-Model Spiking Neural Network</h1>
<blockquote class="blockquote">
<p>Future developments in artificial intelligence will profit from the existence of novel, non-traditional substrates for brain-inspired computing. Neuromorphic computers aim to provide such a substrate that reproduces the brain’s capabilities in terms of adaptive, low-power information processing. We present results from a prototype chip of the BrainScaleS-2 mixed-signal neuromorphic system that adopts a physical-model approach with a 1000-fold acceleration of spiking neural network dynamics relative to biological real time. Using the embedded plasticity processor, we both simulate the Pong arcade video game and implement a local plasticity rule that enables reinforcement learning, allowing the on-chip neural network to learn to play the game. The experiment demonstrates key aspects of the employed approach, such as accelerated and flexible learning, high energy efficiency and resilience to noise.</p>
</blockquote>
<p>https://arxiv.org/abs/1909.11145v1</p>
<p>This paper was presented at ICANN 2019, and published in Lecture Notes in Computer Science. In case it’s unclear what’s going on here: The authors built a small-scale prototype (32 neurons, 32 synapses each) of an apparently <em>analog</em> hardware simulation of a biological learning model of the brain (<a href="https://en.wikipedia.org/wiki/Spike-timing-dependent_plasticity">STDP</a>). They then used it to a) simulate a simplified Pong (on-chip), and b) successfully learn to play using reinforcement learning (again, on-chip). Emulating their own system on an Intel i7-4771 was an order of magnitude slower, so we’re talking about a real improvement. This is an auspicious beginning, and they hint at scaled-up work to come.</p>
<p>I look forward to specialized neuronal hardware. I’m especially interested to hear that they simulated actual neurons to some degree, with spike-timing dependence, rather than the simplified model that I’m used to working with. I expect this means they intend to simulate actual brains at some point. Stay tuned.</p>
</section>
<section id="pre-training-as-batch-meta-reinforcement-learning-with-time" class="level1">
<h1>4. Pre-training as Batch Meta Reinforcement Learning with tiMe</h1>
<blockquote class="blockquote">
<p>Pre-training is transformative in supervised learning: a large network trained with large and existing datasets can be used as an initialization when learning a new task. Such initialization speeds up convergence and leads to higher performance. In this paper, we seek to understand what the formalization for pre-training from only existing and observational data in Reinforcement Learning (RL) is and whether it is possible. We formulate the setting as Batch Meta Reinforcement Learning. We identify MDP mis-identification to be a central challenge and motivate it with theoretical analysis. Combining ideas from Batch RL and Meta RL, we propose tiMe, which learns distillation of multiple value functions and MDP embeddings from only existing data. In challenging control tasks and without fine-tuning on unseen MDPs, tiMe is competitive with state-of-the-art model-free RL method trained with hundreds of thousands of environment interactions.</p>
</blockquote>
<p>https://arxiv.org/abs/1909.11373v1</p>
<p>This paper attempts to bring the benefits of pretraining (on some pre-recorded batch) to reinforcement learning. This is non-trivial, since Q-learning algorithms are known to be unstable on batches produced by “foreign policy” (my phrase).</p>
<blockquote class="blockquote">
<p>The value function diverges if Q fails to accurately estimate the value of <img src="https://latex.codecogs.com/png.latex?%5Cpi(s')"></p>
</blockquote>
<p>This is mitigated by online Q-learning algorithms because the contents of the replay buffer, while produced by an off-policy algorithm, was at least produced through interaction with the environment, and so the distribution of the induced <img src="https://latex.codecogs.com/png.latex?%5Cpi"> doesn’t deviate too much from the distribution in the replay buffer. Even then, this phenomenon is still a source of instability for Q-learning.</p>
<p>In batch learning, the problem is worse. The recorded batch was <em>not</em> produced by our induced policy, and perhaps not even by a <em>single</em> policy. Further, the environment reflected in the batch may not even have been produced by a single Markov decision process.</p>
<p>I’m interested in <em>this</em> paper because the authors apply meta RL to the problem, and claim to achieve good performance on unseen MDPs sampled from the same family as those represented by the training batch. If that’s so, it has positive implications for my own work.</p>
</section>
<section id="reinforcement-learning-with-chromatic-networks" class="level1">
<h1>5. Reinforcement Learning with Chromatic Networks</h1>
<blockquote class="blockquote">
<p>We present a neural architecture search algorithm to construct compact reinforcement learning (RL) policies, by combining ENAS and ES in a highly scalable and intuitive way. By defining the combinatorial search space of NAS to be the set of different edge-partitionings (colorings) into same-weight classes, we represent compact architectures via efficient learned edge-partitionings. For several RL tasks, we manage to learn colorings translating to effective policies parameterized by as few as 17 weight parameters, providing &gt;90% compression over vanilla policies and 6x compression over state-of-the-art compact policies based on Toeplitz matrices, while still maintaining good reward. We believe that our work is one of the first attempts to propose a rigorous approach to training structured neural network architectures for RL problems that are of interest especially in mobile robotics with limited storage and computational resources.</p>
</blockquote>
<p>https://arxiv.org/abs/1907.06511v2</p>
<p>From the introduction:</p>
<blockquote class="blockquote">
<p>The main question we tackle in this paper is the following:</p>
</blockquote>
<blockquote class="blockquote">
<p>Are high dimensional architectures necessary for encoding efficient policies and if not, how compact can they be in in practice?</p>
</blockquote>
<p>More compact achitectures not only take less space, but also produce inferences more quickly and cheaply. This matters to me because my work is often done on cloud computing infrastructure, which incentivizes parsomony. I’m also professionally interested in neural architecture search for multi-task scaling purposes. More on this later, perhaps.</p>
<p>The authors find compact policies by jointly optimizing the RL objective and “the combinatorial nature of the network’s parameter sharing profile”. Inspired by two <a href="https://arxiv.org/abs/1804.02395">other</a> <a href="https://arxiv.org/abs/1906.04358">papers</a>, they reduce the number of distinct weights by <em>sharing</em> a single weight between multiple neuronal connections. The first paper from which their inspiration for this arises used <a href="https://en.wikipedia.org/wiki/Toeplitz_matrix">Toeplitz matrices</a> to represent the neural network, and the second randomly assigns weights (Weight-Agnostic Neural Networks, or WANNs) and then learns the connection topology to maximize an RL goal.</p>
<blockquote class="blockquote">
<p>WANNs replace conceptually simple feedforward networks with general graph topologies using NEAT algorithm providing topological operators to build the network.</p>
</blockquote>
<blockquote class="blockquote">
<p>Our approach is a middle ground, where the topology is still a feedforward neural network, but the weights are partitioned into groups that are being learned in a combinatorial fashion using rein- forcement learning. While <a href="https://arxiv.org/abs/1504.04788">10</a> shares weights randomly via hashing, we learn a good partitioning mechanisms for weight sharing.</p>
</blockquote>
<p>How do they do this?</p>
<blockquote class="blockquote">
<p>We leverage recent advances in the ENAS (Efficient Neural Architecture Search) literature and theory of pointer networks to optimize over the combinatorial component of this objective and state of the art evolution strategies (ES) methods to optimize over the RL objective.</p>
</blockquote>
<blockquote class="blockquote">
<p>Our key observation is that ENAS and ES can naturally be combined in a highly scalable but conceptually simple way.</p>
</blockquote>
<p>Ah. So… <em>how</em> do they do this?</p>
<p>We’ll both just have to read the whole paper. In my light read, I notice this one is so full of interesting insights and pointers to important results from other research that it’s worth our time. Basically though, they alternate between neural architecture search and RL optimization, using their own ENAS variant to optimize a pointer network capable of partitioning weights to be shared.</p>
</section>
<section id="dynamical-distance-learning-for-semi-supervised-and-unsupervised-skill-discovery" class="level1">
<h1>6. Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery</h1>
<blockquote class="blockquote">
<p>Reinforcement learning requires manual specification of a reward function to learn a task. While in principle this reward function only needs to specify the task goal, in practice reinforcement learning can be very time-consuming or even infeasible unless the reward function is shaped so as to provide a smooth gradient towards a successful outcome. This shaping is difficult to specify by hand, particularly when the task is learned from raw observations, such as images. In this paper, we study how we can automatically learn dynamical distances: a measure of the expected number of time steps to reach a given goal state from any other state. These dynamical distances can be used to provide well-shaped reward functions for reaching new goals, making it possible to learn complex tasks efficiently. We show that dynamical distances can be used in a semi-supervised regime, where unsupervised interaction with the environment is used to learn the dynamical distances, while a small amount of preference supervision is used to determine the task goal, without any manually engineered reward function or goal examples. We evaluate our method both on a real-world robot and in simulation. We show that our method can learn to turn a valve with a real-world 9-DoF hand, using raw image observations and just ten preference labels, without any other supervision. Videos of the learned skills can be found on the project website: <a href="https://sites.google.com/view/dynamical-distance-learning">https://sites.google.com/view/dynamical-distance-learning</a></p>
</blockquote>
<p>https://arxiv.org/abs/1907.08225v2</p>
<p>I picked up this paper partly because reward shaping is currently of professional interest to me, but also because I’m watching Haarnoja for his work on distributional RL.</p>
<p>This paper is about making reward-shaping easier by learning a more direct distance measure for the purpose. In general, if you know your distance from a goal, there are many optimization methods available to you for reducing that distance and achieving your goal. The better this distance measure, the smoother the landscape, and the more quickly you arrive.</p>
<p>The semi-supervised way_in which they approach this problem also strikes me as relevant to AI safety, a topic of personal interest. With the unsupervised training of their dynamical distance measure, they add “a small amount of preference supervision” to set the task goal, and this results in its achievement. A manually-specified reward function is very dangerous, and I’m interested in any novel methods that avoid their direct use (or, more to the point, I’m interested in methods of motivating AIs which more directly align holistic human flourishing with the AI’s objectives).</p>
<p>For a quick overview, don’t miss the link they posted at the end of the abstract.</p>
</section>
<section id="model-imitation-for-model-based-reinforcement-learning" class="level1">
<h1>7. Model Imitation for Model-Based Reinforcement Learning</h1>
<blockquote class="blockquote">
<p>Model-based reinforcement learning (MBRL) aims to learn a dynamic model to reduce the number of interactions with real-world environments. However, due to estimation error, rollouts in the learned model, especially those of long horizon, fail to match the ones in real-world environments. This mismatching has seriously impacted the sample complexity of MBRL. The phenomenon can be attributed to the fact that previous works employ supervised learning to learn the one-step transition models, which has inherent difficulty ensuring the matching of distributions from multi-step rollouts. Based on the claim, we propose to learn the synthesized model by matching the distributions of multi-step rollouts sampled from the synthesized model and the real ones via WGAN. We theoretically show that matching the two can minimize the difference of cumulative rewards between the real transition and the learned one. Our experiments also show that the proposed model imitation method outperforms the state-of-the-art in terms of sample complexity and average return.</p>
</blockquote>
<p>https://arxiv.org/abs/1909.11821v1</p>
<p>AGI seems likely to be model-based, rather than model-free. I think this because I (an AGI) personally reuse my own models all the time, frequently attempt near-transfer to solve some novel problem. So anything that claims progress on model-based learning is at least worth a look to me.</p>
<p>Earlier I blogged about <a href="http://localhost:8000/articles/2019/Jul/28/efficient-exploration-with-self-imitation-learning.html">Efficient Exploration with Self-Imitation Learning via Trajectory-Conditioned Policy</a>, and looking at it now, I’m surprised I only <em>alluded</em> to their use of Transformers. In that paper, they learn to imitate a past trajectory by mimicking the trajectory distribution, conditioned on a past trajectory. <em>This</em> paper wants to create a model of the environment that similarly mimics its distribution, but they use Wasserstein GANs (WGANs) instead. GANs have been wildly successful in generative image models, and WGANs are an especially promising variant. I’ve been keeping an eye out for papers that use GANs in areas outside computer vision.</p>
<p>If you know what a WGAN is and you understand that the authors are trying to get a WGAN to mimic the environment’s bounded trajectory segment transition distribution, then you can imagine what they’re doing. They also provide a theoretical bound for the expected distributional error.</p>
<p>I’ll mention in closing that in their experiments, they also end up doing better than most other methods, using 50% fewer samples. That it works at all suggests to me that the model is sufficiently accurate to take notice. GANs are interesting.</p>
</section>
<section id="zermelos-problem-optimal-point-to-point-navigation-in-2d-turbulent-flows-using-reinforcement-learning" class="level1">
<h1>8. Zermelo’s problem: Optimal point-to-point navigation in 2D turbulent flows using Reinforcement Learning</h1>
<blockquote class="blockquote">
<p>To find the path that minimizes the time to navigate between two given points in a fluid flow is known as Zermelo’s problem. Here, we investigate it by using a Reinforcement Learning (RL) approach for the case of a vessel which has a slip velocity with fixed intensity, Vs , but variable direction and navigating in a 2D turbulent sea. We show that an Actor-Critic RL algorithm is able to find quasi-optimal solutions for both time-independent and chaotically evolving flow configurations. For the frozen case, we also compared the results with strategies obtained analytically from continuous Optimal Navigation (ON) protocols. We show that for our application, ON solutions are unstable for the typical duration of the navigation process, and are therefore not useful in practice. On the other hand, RL solutions are much more robust with respect to small changes in the initial conditions and to external noise, even when V s is much smaller than the maximum flow velocity. Furthermore, we show how the RL approach is able to take advantage of the flow properties in order to reach the target, especially when the steering speed is small.</p>
</blockquote>
<p>https://arxiv.org/abs/1907.08591v2</p>
<p>This paper is a pure personal indulgence. I read James Gleick’s <a href="https://www.amazon.com/gp/product/0143113453">Chaos</a> and got interested in dynamical systems theory. I don’t know much, but I do know turbulent flows are a pain to predict, which I assume would mean they’re a pain to navigate within. I’ve also heard that <a href="https://link.springer.com/article/10.1007/BF02312352">neural networks do pretty surprisingly well at predicting chaotic dynamics</a>, so I’m interested to see it “applied”. The paper brings up several other examples of successful neural navigation and prediction:</p>
<blockquote class="blockquote">
<p>Promising results have been obtained when applying RL algorithms to similar problems, such as the training of smart inertial particles or swimming particles navigat- ing intense vortex regions [31], Taylor Green flows [32] and ABC flows [33]. RL has also been successfully imple- mented to reproduce schooling of fishes [34, 35], soaring of birds in a turbulent environments [36, 37] and in many other applications [38–40]. Similarly, in the recent years, artificial intelligence techniques are establishing them- selves as new data driven models for fluid mechanics in general [41–46].</p>
</blockquote>
<p>I can’t state their results better than they can, so here you go:</p>
<blockquote class="blockquote">
<p>In this paper, we show that for the case of vessels that have a slip velocity with fixed intensity but variable di- rection, RL can find a set of quasi-optimal paths to efficiently navigate the flow. Moreover, RL, unlike ON, can provide a set of highly stable solutions, which are insensitive to small disturbances in the initial condition and successful even when the slip velocity is much smaller than the guiding flow. We also show how the RL protocol is able to take advantage of different features of the underlying flow in order to achieve its task, indicating that the information it learns is non-trivial.</p>
</blockquote>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li>Surprisingly, even this list of eight do not cover <em>all</em> of the papers that sounded interesting to me. It was <em>quite</em> a good week for announcements on the arXiv.</li>
<li>This is the format I originally had in mind for arXiv highlights, but since the abstracts tend to invite questions that I can’t answer without at least skimming the paper, I ended up reading them more thoroughly. With this format, I can cover more ground, but less deeply.</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/comments-on-eight-abstracts/</guid>
  <pubDate>Sun, 06 Oct 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/comments-on-eight-abstracts.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>Active Perception in Adversarial Scenarios</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/active-perception-in-adversarial-scenarios/</link>
  <description><![CDATA[ 





<section id="this-week" class="level1">
<h1>This week</h1>
<p>This week’s paper is <a href="https://arxiv.org/abs/1902.05644v1">Active Perception in Adversarial Scenarios using Maximum Entropy Deep Reinforcement Learning</a>. The idea is that an agent interacting with another agent can learn to assess the threat it may pose. It does this by actively testing the opponent agent’s behavior, and does not assume the opponent’s behavior remains stationary. It uses Bayesian filtering to update its belief about the disposition of the opponent, and that’s why this paper caught my eye. I’m on a Bayesian kick lately.</p>
<blockquote class="blockquote">
<p>To summarize, the contribution here is the development of a scalable robust active perception method in scenarios where a potential adversary opponent could be actively hostile to the intent recognition activity, which extends and outperforms the POMDP methods.</p>
</blockquote>
<p>I’m a bit short on time this week, so I apologize for the amount of jargon and the unusually high level of confusion.</p>
</section>
<section id="problem-setup" class="level1">
<h1>Problem setup</h1>
<blockquote class="blockquote">
<p>We model the active perception problem as a planning problem, defined by the tuple <img src="https://latex.codecogs.com/png.latex?%5Clangle%20S,A%5Ea,A%5Eo,T,O,R,b_0,%5Cgamma%20%5Crangle">, where <img src="https://latex.codecogs.com/png.latex?S=%5Clangle%20S%5Eo,S%5Ep%20%5Crangle"> is the state of the world, consisting of the set of observable states <img src="https://latex.codecogs.com/png.latex?S%5Eo"> and the set of partially observable states <img src="https://latex.codecogs.com/png.latex?S%5Ep">; <img src="https://latex.codecogs.com/png.latex?A%5Ea"> is the set of actions of the autonomous agent; <img src="https://latex.codecogs.com/png.latex?A%5Eo"> is the set of actions of the opponent; we further assume that regardless of the intention, the opponent has the same set of observable actions. Otherwise, an intention is easily identifiable once an action that is uniquely corresponding to that type of intention is observed. $T:S A^a A^o _S $ is the transition probability, where <img src="https://latex.codecogs.com/png.latex?%5CDelta_%7B%5Cbullet%7D"> denotes the space of probability distribution over the space <img src="https://latex.codecogs.com/png.latex?%5Cbullet">. <img src="https://latex.codecogs.com/png.latex?O:%20S%20%5Ctimes%20A%5Ea%20%5Crightarrow%20%5CDelta_%7BA%5Eo%7D"> is the observation probability; <img src="https://latex.codecogs.com/png.latex?R:%20S%20%5Ctimes%20A%5Ea%20%5Ctimes%20A%5Eo%20%5Crightarrow%20%20%5Cmathbb%7BR%7D"> is the reward function; <img src="https://latex.codecogs.com/png.latex?b_0"> is the prior probability of the opponent being an adversary; and <img src="https://latex.codecogs.com/png.latex?%5Cgamma"> is the discount factor.</p>
</blockquote>
<p>Further, the opponent is assumed to be either neutral (merely self-interested, in a known way) or hostile (goal-directed, as defined by a known MDP), with bounded rationality, (it may not be able to take the optimal action) and it is likely to behave deceptively.</p>
<p>Notice that the actual behavior of the opponent is known if its disposition is known, which to my mind may or may not be a reasonable assumption, depending on the setting. Since I’ve had AI safety on the brain lately, it strikes me as <em>unrealistic</em> in a situation where your opponent is smarter than you are. It may be more realistic in settings where everyone has the same goal and it’s relatively clear how anyway would try to achieve it if they didn’t have to deal with other agents.</p>
<p>The authors’ adversarial model is interesting. (<img src="https://latex.codecogs.com/png.latex?%5Clambda"> is the parameter to <img src="https://latex.codecogs.com/png.latex?%5Cpi%5Eo"> that specifies whether the agent is neutral: <img src="https://latex.codecogs.com/png.latex?%5Clambda=0">, or adversarial: <img src="https://latex.codecogs.com/png.latex?%5Clambda=1">):</p>
<blockquote class="blockquote">
<p>We use the following equation to model an adversarial agent’s policy <img src="https://latex.codecogs.com/png.latex?%5Cpi%5Eo">: <img src="https://latex.codecogs.com/png.latex?%0A%5Cbegin%7Balign%7D%0A%5Cpi%5Eo(a%5Eo_t%7Cs_t,%5Clambda=1;%5Calpha,%5Cbeta)=%20&amp;%20%5Ctext%7Bargmin%7D_%7B%5Cpi%20%5Cin%20%5CDelta%7D%20%5C%7B%5Cmathbb%7BKL%7D(%5Cpi%7C%5Cpi%5E%7B%5Ctext%7BMDP%7D%7D_%7B%5Calpha%7D)%5C%5C%0A&amp;%20+%5Cbeta%20%5Cmathbb%7BKL%7D(%5Cpi%7C%5Cpi%5Eo(%5Ccdot%7Cs_t,%5Clambda=0))%20%5C%7D%0A%5Cpi%5E%7B%5Ctext%7BMDP%7D%7D_%7B%5Calpha%7D(a_t%5Eo%7Cs_t,%5Clambda=1)=e%5E%7B%5Calpha%20Q(s_t,a_t%5Eo)%7D/Z(s_t)%0A%5Cend%7Balign%7D%0A"></p>
</blockquote>
<p>The thing to take away from this is that both rationality and deception are tunable parameters. The rationality of the opponent is controlled by the temperature parameter <img src="https://latex.codecogs.com/png.latex?%5Calpha">, by adjusting how well the opponent makes use of the optimal Q function. The degree to which the opponent is <em>deceptive</em> is controlled by <img src="https://latex.codecogs.com/png.latex?%5Cbeta">, which adjusts how much the KL-divergence of the existing policy from the <em>neutral</em> policy affects the opponent’s search for an optimal strategy.</p>
</section>
<section id="bayesian-filtering" class="level1">
<h1>Bayesian filtering</h1>
<blockquote class="blockquote">
<p>We maintain a belief <img src="https://latex.codecogs.com/png.latex?b_t(%5Clambda)"> over the hidden variable by Bayesian filtering.</p>
</blockquote>
<p>As I mentioned, I’m rather short on time today, so I must apologize again for not actually spending the time to explain this. For now, suffice it to say that the opponent is either neutral (<img src="https://latex.codecogs.com/png.latex?%5Clambda=0">) or hostile (<img src="https://latex.codecogs.com/png.latex?%5Clambda=1">), and how your agent reacts to it depends very much on which one of those it believes it is playing against. Bayesian filtering will allow it to make the most of the evidence available, so it can use its best guess as it trains.</p>
<blockquote class="blockquote">
<p>We define a hybrid belief-state dependent reward to balance exploration and safety <img src="https://latex.codecogs.com/png.latex?%5Cbegin%7Bequation%7D%0A%5Cbegin%7Baligned%7D%0Ar(b_t,s_t,a%5Ea_t)&amp;=-H(b_t)+r(s_t,a%5Ea_t)%5C%5C%0A&amp;=b%5Clog%20b+(1-b)%5Clog(1-b)+r(s_t,a%5Ea_t),%0A%5Cend%7Baligned%7D%0A%5Clabel%7Beq6%7D%0A%5Cend%7Bequation%7D"> where we use the shorthand <img src="https://latex.codecogs.com/png.latex?b"> to denote <img src="https://latex.codecogs.com/png.latex?b_t(%5Clambda=1)">, the belief that the opponent is an adversary; and <img src="https://latex.codecogs.com/png.latex?r(s_t,a%5Ea_t)"> is the state dependent reward.</p>
</blockquote>
<blockquote class="blockquote">
<p>This reward balances exploration behavior and safety. The negative entropy reward <img src="https://latex.codecogs.com/png.latex?-H(b_t)"> can be interpreted as maximizing the expected logarithm of true positive rate (TPR) and true negative rate (TNR). The state-dependent reward <img src="https://latex.codecogs.com/png.latex?r(s_t,a%5Ea_t)"> depends both on the observable state and the partially observable intent state <img src="https://latex.codecogs.com/png.latex?%5Clambda">, as well as the action of the autonomous agent. This reward is used to ensure safety. For instance, some actions could be dangerous to the neutral [opponent], which are discouraged by a large negative reward.</p>
</blockquote>
<p>Our agent is trained using <a href="https://arxiv.org/abs/1702.08165">Soft-Q Learning</a> while values of <img src="https://latex.codecogs.com/png.latex?%5Clambda"> are varied, with corresponding opponent behavior. Interestingly, in the case study section the authors mention that the actual adversary models were not always provided in the learning phase.</p>
<blockquote class="blockquote">
<p>The active perception agent has to identify the hidden intent while bein grobust to this model uncertainty, which is challenging.</p>
</blockquote>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li>I admit to being a bit confused by this paper. The authors claim to do Bayesian filtering, but it’s not an explicit feature of the algorithm. In fact, they seem to be sampling <img src="https://latex.codecogs.com/png.latex?%5Clambda"> for use in training by using only <img src="https://latex.codecogs.com/png.latex?b_0">, their prior probability for their belief state. Perhaps it’s a typo.</li>
<li>They also seem to claim that the two models of the opponent behavior must be known, but then they mention they’re not available during the learning phase in their case study. Drop me a line if this makes sense to you.</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/active-perception-in-adversarial-scenarios/</guid>
  <pubDate>Sun, 22 Sep 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/active-perception-in-adversarial-scenarios.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>Discovery of Useful Questions as Auxiliary Tasks</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/discovery-of-useful-questions-as-auxiliary-tasks/</link>
  <description><![CDATA[ 





<p>In case you’re wondering what happened to your feed reader this week: We’ve decided to retitle all of the arXiv highlights posts to be more attractive. We promise not to do this often, but it seemed like a good time to do it while we’re inconveniencing very few people.</p>
<section id="this-week" class="level1">
<h1>This week</h1>
<p>This week’s paper is <a href="https://arxiv.org/abs/1909.04607v1">Discovery of Useful Questions as Auxiliary Tasks</a> from the University of Michigan and DeepMind. It was accepted to NeurIPS 2019 (which I rather hope I’ll be attending). The paper contains a very exciting concept that strikes at the heart of human learning: We learn not only by noticing statistical correlations and inferring concepts, but by actively seeking the answers to helpful questions that occur to us as we navigate the world. That’s also much of what science is about: increasing your understanding of the world by choosing particularly good questions to ask.</p>
</section>
<section id="useful-questions-as-an-auxiliary-task" class="level1">
<h1>Useful questions as an auxiliary task</h1>
<p>The authors formulate the problem as a reinforcement learning problem with a main task you’d like to accomplished, augmented with auxiliary tasks generated by the system itself to aid in representation learning, and ultimately to accomplish the main task more efficiently. I’ve mentioned before that this is of professional interest to me.</p>
<p>In this paper the questions are represented as “general value functions” (GVFs), “a fairly rich form of knowledge representation”, because</p>
<blockquote class="blockquote">
<p>GVF-based auxiliary tasks have been shown in previous work to improve the sampling efficiency of reinforcement learning agents engaged in learning some complex task…. It was then shown that by combining gradients from learning the auxiliary GVFs with the updates from the main task, it was possible to accelerate representation learning and improve performance. It fell, however, onto the algorithm designer to design questions that were useful for the specific task.</p>
</blockquote>
<p>The main insight in this paper is that the gradients induced while learning the main task contain information about what questions would aid in learning a helpful representation.</p>
<blockquote class="blockquote">
<p>The main idea is to use meta-gradient RL to discover the questions so that answering them maximises the usefulness of the induced representation on the main task.</p>
</blockquote>
</section>
<section id="auxiliary-tasks" class="level1">
<h1>Auxiliary tasks</h1>
<p>Why should learning something other than the main task help? It teaches composable fundamentals relevant to the task so that the neural network doesn’t have to learn everything from scratch all at once. The kinds of auxiliary tasks we’re talking about here are things like controlling pixel intensities and feature activations. Other examples mentioned in the paper are auxiliary tasks where the agent needed to learn to measure depth, loop-closures (e.g., the letter “C” is not closed, but the letter “O” is), observation reconstruction (which, as an aside, can be used in the construction of intrinsically-motivated, “curious” agents), reward prediction, etc. When agents were required to learn each of these tasks simultaneously with learning their own main tasks, they learned more efficiently than when they were required to learn their main task alone.</p>
<p>But, as we just discussed, each of these examples (see the paper for more) and were hand-crafted. The agents themselves did not attempt to add to their tasks, and careful hand-tuning was required to get the observed improvements.</p>
</section>
<section id="meta-learning" class="level1">
<h1>Meta-learning</h1>
<blockquote class="blockquote">
<p>A meta-learner progressively improves the learning process of a learner that is attempting to solve some task.</p>
</blockquote>
<p>I can hardly overstate how useful this is. In my own work, we aren’t done as soon as we’ve trained a neural network to perform well on a single task. There is an entire host of related tasks on which we’ll need to retrain it in the future. Our work involves training an agent to control the behavior of some software, which is not fixed. If our agent cannot be quickly retrained on other software (perhaps out of our direct control), then it becomes much more expensive and difficult to maintain.</p>
<p>This paper mentions previous work in learning better initializations for a given task, learning to explore, unsupervised learning to develop a good or compact representation, few-shot model adaptation, and learning to improve the optimizers.</p>
</section>
<section id="the-discovery-of-useful-questions" class="level1">
<h1>The discovery of useful questions</h1>
<p>This is Figure 1 of our paper, depicting the architecture that discovers and uses useful questions. It consists of two neural networks, a main task &amp; answer network parametrized by <img src="https://latex.codecogs.com/png.latex?%5Ctheta">, and a question network parametrized by <img src="https://latex.codecogs.com/png.latex?%5Ceta">. The main task &amp; answer network takes the last <img src="https://latex.codecogs.com/png.latex?i"> observations <img src="https://latex.codecogs.com/png.latex?o_%7Bt-i+1:t%7D"> in and produces two categories of output: a) decisions from the policy <img src="https://latex.codecogs.com/png.latex?%5Cpi_t"> and b) answers to the “useful questions” <img src="https://latex.codecogs.com/png.latex?y_t">. The question network takes <img src="https://latex.codecogs.com/png.latex?j"> <em>future</em> observations <img src="https://latex.codecogs.com/png.latex?o_%7Bt+1:t+j%7D">, and produces two outputs: a) <em>cumulants</em> <img src="https://latex.codecogs.com/png.latex?u_t">, and b) discounts <img src="https://latex.codecogs.com/png.latex?%5Cgamma_t">. Cumulants (a term from the GVF literature) are described as scalar functions of the state, the sum of which must be maximized. To me, this just sounds like an obstruse way to say “other loss function”, which makes sense because these are what are describing our auxiliary goals.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://computable.ai/static/images/useful_questions_figure1.png" class="img-fluid figure-img"></p>
<figcaption>Auxiliary Question Discovery Arch</figcaption>
</figure>
</div>
<p>Lest you think this method requires time travel, fear not. We can see <img src="https://latex.codecogs.com/png.latex?j"> steps into the future using the time machine of Waiting, which is ok because it only happens during training.</p>
<p>As the authors explain, previous work with auxiliary tasks would have only had the main task &amp; answer network on the left, because the cumulants and discounts were hand-crafted. The question network on the right, and its effective use, is the main contribution of this paper. The <em>number</em> of “other loss functions” is still fixed, but the components of the actual functions that compute them (cumulants and discounts) are represented by an <img src="https://latex.codecogs.com/png.latex?%5Ceta">-parametrized neural network that is itself trained <em>on the gradients of the <img src="https://latex.codecogs.com/png.latex?%5Ctheta">-parametrized main task and answer network</em>.</p>
<p>In the researcher’s own words:</p>
<blockquote class="blockquote">
<p>In their most abstract form, reinforcement learning algorithms can be described by an update procedure <img src="https://latex.codecogs.com/png.latex?%5CDelta%20%5Ctheta_t"> that modifies, on each step <img src="https://latex.codecogs.com/png.latex?t">, the agent’s parameters <img src="https://latex.codecogs.com/png.latex?%5Ctheta_t">. The central idea of meta-gradient RL is to parameterise the update <img src="https://latex.codecogs.com/png.latex?%5CDelta%20%5Ctheta_t(%5Ceta)"> by meta-parameters <img src="https://latex.codecogs.com/png.latex?%5Ceta">. We may then consider the consequences of changing <img src="https://latex.codecogs.com/png.latex?%5Ceta"> on the <img src="https://latex.codecogs.com/png.latex?%5Ceta">-parameterised update rule by measuring the subsequent performance of the agent, in terms of a “meta-loss” function <img src="https://latex.codecogs.com/png.latex?m(%5Ctheta_%7Bt+k%7D)">. Such meta-loss may be evaluated after one update (myopic) or <img src="https://latex.codecogs.com/png.latex?k%20%3E%201"> updates (non-myopic). The meta-gradient is then, by the chain rule, <img src="https://latex.codecogs.com/png.latex?%5Cbegin%7Balign%7D%0A%7B%5Cpartial%20m(%5Ctheta_%7Bt+k%7D)%7D%20%5Cover%20%7B%5Cpartial%5Ceta%7D%20&amp;=%20%7B%5Cpartial%20m(%5Ctheta_%7Bt+k%7D)%20%5Cover%20%5Cpartial%5Ctheta_%7Bt+k%7D%7D%20%7B%5Cpartial%5Ctheta_%7Bt+k%7D%20%5Cover%20%5Cpartial%5Ceta%7D.%5Clabel%7Beqn:no_approx%7D%0A%5Cend%7Balign%7D"></p>
</blockquote>
<p>The actual computation of this is challenging, because changing <img src="https://latex.codecogs.com/png.latex?%5Ceta"> affects updates to <img src="https://latex.codecogs.com/png.latex?%5Ctheta"> on <em>all future timesteps</em>. This is the reason training the question network requires looking <img src="https://latex.codecogs.com/png.latex?j"> steps “into the future”. Holding <img src="https://latex.codecogs.com/png.latex?%5Ceta"> fixed, they compute <img src="https://latex.codecogs.com/png.latex?%5Ctheta_t%20%5Crightarrow%20...%20%5Crightarrow%20%5Ctheta_%7Bt+j%7D">, in order to finally compute the meta-loss evaluation <img src="https://latex.codecogs.com/png.latex?m(%5Ctheta_%7Bt+j%7D)">.</p>
<p>The algorithm then alternates between normal RL training of the main task &amp; answer network, and meta-gradient training of the question network to produce and use questions that maximize the performance of the agent on the original task. It is a very general solution, and empirically outperforms hand-designed auxiliary tasks in many cases.</p>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li>The authors themselves note that their algorithm augments an <em>on-policy</em> reinforcement learning algorithm, and I look forward to their promised future work adapting these techniques to an off-policy setting.</li>
<li>I notice I take detours from the main article purposes to write about areas of RL that I want to remember to investigate further in the future (e.g., auxiliary task in general, and meta-learning in general). That’s a good habit, though I’ll need to remember to cultivate it without seeming too distracted.</li>
<li>This paper mentions that Xu et al.&nbsp;in 2018 tried learning the discount factor <img src="https://latex.codecogs.com/png.latex?%5Cgamma"> and the bootstrapping factor <img src="https://latex.codecogs.com/png.latex?%5Clambda"> (using meta-gradients), which is an idea I had myself (a year later). Apparently this substantially improved performance on the Atari domain, so I feel vindicated.</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/discovery-of-useful-questions-as-auxiliary-tasks/</guid>
  <pubDate>Sun, 15 Sep 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/discovery-of-useful-questions-as-auxiliary-tasks.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>Deep Reinforcement Learning without Catastrophic Forgetting</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/deep-reinforcement-learning-without-catastrophic-forgetting/</link>
  <description><![CDATA[ 





<p>Apologies for missing a week. Today’s post is on last-week’s paper, and I’m going to skip this week to get back on track. Also experimenting with the format some more to keep things sustainable given my wildly variable weekend free time. If you have thoughts about this, please leave us a comment!</p>
<section id="this-week" class="level1">
<h1>This week</h1>
<p>This (last) week’s paper is <a href="https://arxiv.org/abs/1812.02464">Pseudo-Rehearsal: Achieving Deep Reinforcement Learning without Catastrophic Forgetting</a>. I’m interested for reasons both professional and personal.</p>
<p>First, I have this problem. Our recent (successful) work has gotten neural nets to do some very interesting things, but expanding will require continuous training in production. This makes catastrophic forgetting (CF) a very real problem, since most of the DRL research assumes you’re training your agent on a single task, and then enjoying it in inference mode forever after.</p>
<p>Second, I’m interested because I’ve got a little son, (the source of the variability in my weekend free time) and I often see him learn something mind-bogglingly fast, and then cement it over the course of a couple days. Pseudo-rehearsal is biologically plausible, and I’m interested in intelligence in its own right.</p>
</section>
<section id="catastrophic-forgetting-and-pseudo-rehearsal" class="level1">
<h1>Catastrophic Forgetting and Pseudo-rehearsal</h1>
<p>An agent trained on one task can learn to accomplish that task. If that same agent is then moved to another task, it will learn that other task, but often at the expense of “catastrophically forgetting” the neural net weights learned for the previous task. Several solutions have been proposed, (which are cited in today’s paper, and I’ll likely be reading them) but most are likely <em>not</em> what humans and animals do.</p>
<blockquote class="blockquote">
<p>Researchers have proposed extensions to this method such as utilising previous examples’ gradients during learning, picking a subset of previous samples which best represents the population and using a variational auto-encoder to compress stored items. Such rehearsal methods are cognitively implausible and therefore, do not shine light on how mammal brains might efficiently solve the CF problem.</p>
</blockquote>
<p>Pseudo-rehearsal trains a generative model (a GAN) to produce examples from all previous tasks, and uses this to implicitly rehearse foregoing data. Today’s paper employes this scheme and a few other tricks to build a system capable of learning multiple tasks.</p>
</section>
<section id="the-repr-model" class="level1">
<h1>The RePR model</h1>
<p>The researchers dub their method RePR, and it works like this: They build short- and long-term memory systems, and transferring learned behaviors from short- to long-term memory while rehearsing past behavior in long-term memory.</p>
<p>The STM system:</p>
<blockquote class="blockquote">
<p>The first part of our model is the short- term memory (STM) system, which serves a similar function to the hippocampus and is used to learn the current task. The STM system contains two components, a DQN that learns the current task and an experience replay containing data only from the current task.</p>
</blockquote>
<p>The LTM system:</p>
<blockquote class="blockquote">
<p>The second part is the long-term memory (LTM) system, which serves a similar function to the cortex. The LTM system also has two components, a DQN containing knowledge of all tasks learnt and a GAN which can generate sequences representative of these tasks.</p>
</blockquote>
<p>They then do periodic consolidation:</p>
<blockquote class="blockquote">
<p>During consolidation, the LTM retains previous knowledge through pseudo-rehearsal, while being taught by the STM how to respond on the current task. All of the networks’ architectures and training parameters used throughout our experiments can be found in the appendices. Transferring knowledge between these two systems is achieved through knowledge distillation, where a student network is optimised so that it outputs similar values to a teacher network.</p>
</blockquote>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li>This sounds brilliant, and analogous to what mammals do. I’m eager to experiment with it, and to introspect and ponder how my own brain learns, with this new model in mind.</li>
<li>I wonder very much what we do in sleep. <a href="https://computable.ai/articles/2019/Mar/10/boltzmann-machines-differentiation-work.html">As I’ve mentioned before</a>, I’m quite attracted to the model described in <a href="https://theneural.wordpress.com/2011/07/08/the-miracle-of-the-boltzmann-machine/">The Miracle of the Boltzmann Machine</a>, but off-hand, I don’t know how to reconcile that model with the concept of nightly rehearsal of the day’s activities. Perhaps the brain is doing <em>two</em> things during sleep? Ockam’s razor impells me to think again.</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/deep-reinforcement-learning-without-catastrophic-forgetting/</guid>
  <pubDate>Mon, 09 Sep 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/deep-reinforcement-learning-without-catastrophic-forgetting.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>Reward tampering</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/reward-tampering/</link>
  <description><![CDATA[ 





<section id="this-week" class="level1">
<h1>This week</h1>
<p>This week I just want to pull the list of reward tampering methods from <a href="https://arxiv.org/abs/1908.04734">Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective</a> to promote awareness of this problem. The paper is interesting for several other reasons as well, and I commend it to you:</p>
<blockquote class="blockquote">
<p>Can an arbitrarily intelligent reinforcement learning agent be kept under control by a human user? Or do agents with sufficient intelligence inevitably find ways to shortcut their reward signal? This question impacts how far reinforcement learning can be scaled, and whether alternative paradigms must be developed in order to build safe artificial general intelligence.</p>
</blockquote>
</section>
<section id="reward-tampering" class="level1">
<h1>Reward tampering</h1>
<p>I’ve heard it said that no agent will ever become more intelligent than it takes to edit its own reward function, giving itself a simpler task. This paper treats such problems seriously, with some encouraging results.</p>
<blockquote class="blockquote">
<p>From an AI safety perspective, we must bear in mind that in any practically implemented system, agent reward may not coincide with user utility. In other words, the agent may have found a way to obtain reward without doing the task. This is sometimes called reward hacking or reward corruption. We distinguish between a few different types of reward hacking.</p>
</blockquote>
<section id="reward-gaming-vs.-reward-tampering" class="level2">
<h2 class="anchored" data-anchor-id="reward-gaming-vs.-reward-tampering">Reward gaming vs.&nbsp;reward tampering</h2>
<p>The authors make a distinction between <em>reward gaming</em>, where the agent exploits a misspecification of the process that determines the rewards, and <em>reward tampering</em>, where the agent actually modifies that process. This paper is focused on the latter.</p>
<p>They then subdivide reward tampering into three subcategories, according to whether the agent has tampered with the function itself, the feedback that trains the reward function, or the input to the reward function.</p>
</section>
<section id="hacking-the-reward-function-section-3" class="level2">
<h2 class="anchored" data-anchor-id="hacking-the-reward-function-section-3">Hacking the reward function: Section 3</h2>
<blockquote class="blockquote">
<p>First, regardless of whether the reward is chosen by a computer program, a human, or both, a sufficiently capable, real-world agent may find a way to tamper with the decision. The agent may for example hack the computer program that determines the reward. Such a strategy may bring high agent reward and low user utility. This reward function tampering problem will be explored in Section 3.</p>
<p>Fortunately, there are modifications of the RL objective that remove the agent’s incentiveto tamper with the reward function.</p>
</blockquote>
<p>In Section 3 the authors formalize the problem, and propose two reward variants that disincentivize tampering.</p>
</section>
<section id="manipulating-the-feedback-mechanism-section-4" class="level2">
<h2 class="anchored" data-anchor-id="manipulating-the-feedback-mechanism-section-4">Manipulating the feedback mechanism: Section 4</h2>
<blockquote class="blockquote">
<p>The related problem of reward gaming can occur even if the agent never tamperswith the reward function. A promising way to mitigate the reward gaming problem isto let the user continuously give feedback to update the reward function, using online reward-modeling. Whenever the agent finds a strategy with high agent reward but low user utility, the user can give feedback that dissuades the agent from continuing the behavior. However, a worry with online reward modeling is that the agent may influence the feedback. For example, the agent may prevent the user from giving feedback while continuing to exploit a misspecified reward function, or manipulate the user to give feedback that boosts agent reward but not user utility. This feedback tampering problem and its solutions will be the focus of Section 4.</p>
</blockquote>
<p>Section 4 proposes several potential modifications to disincentivize or directly prevent feedback manipulation, ultimately with the recommendation that they be combined in an ensemble.</p>
</section>
<section id="input-tampering-section-5" class="level2">
<h2 class="anchored" data-anchor-id="input-tampering-section-5">Input tampering: Section 5</h2>
<blockquote class="blockquote">
<p>Finally, the agent may tamper with the input to the reward function, so-called RF-input tampering, for example by gluing a picture in front of its camera to fool the reward function that the task has been completed. This problem and its potential solution will be the focus of Section 5.</p>
</blockquote>
<p>Very interestingly, Section 5 argues that model-based methods avoid the input tampering problem.</p>
</section>
</section>
<section id="results-summary" class="level1">
<h1>Results summary</h1>
<blockquote class="blockquote">
<p>One way to prevent the agent from tampering with the reward function is to isolate or encrypt the reward function, and in other ways trying to physically prevent the agent from reward tampering. However, we do not expect such solutions to scale indefinitely with our agent’s capabilities, as a sufficiently capable agent may find ways around most defenses. Instead, we have argued for design principles that prevent reward tampering incentives, while still keeping agents motivated to complete the original task. Indeed, for each type of reward tampering possibility, we described one or more design principles for removing the agent’s incentive to use it. The design principles can be combined into agent designs with no reward tampering incentive at all.</p>
<p>An important next step is to turn the design principles into practical and scalable RL algorithms, and to verify that they do the right thing in setups where various types of reward tampering are possible. With time, we hope that these design principles will evolve into a set of best practices for how to build capable RL agents without reward tampering incentives. We also hope that the use of causal influence diagrams that we have pioneered in this paper will contribute to a deeper understanding of many other AI safety problems and help generate new solutions.</p>
</blockquote>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li>I look forward to reading this paper more thoroughly, both because I understand this problem of disincentivising reward hacking is <em>hard</em>, and because Causal Influence Diagrams sound interesting and generally useful.</li>
<li>AI safety is important, and I rather hope that awareness of some ways your agents could cheat will help to prevent such errors from leaking out into the world before they are caught.</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/reward-tampering/</guid>
  <pubDate>Sun, 25 Aug 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/reward-tampering.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>DRL Not Superhuman on Atari</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/drl-not-superhuman-on-atari/</link>
  <description><![CDATA[ 





<section id="this-week" class="level1">
<h1>This week</h1>
<p>Just a sketch this week, calling your attention to <a href="https://arxiv.org/abs/1908.04683v1">Is Deep Reinforcement Learning Really Superhuman on Atari?</a>, which concludes not only that DRL is worse than the best humans on most Atari games, but by a <em>wide</em> margin.</p>
</section>
<section id="drl-isnt-superhuman-on-atari-yet" class="level1">
<h1>DRL <em>isn’t</em> superhuman on Atari yet</h1>
<p>Wait, what? I was quite skeptical of this claim. Mnih et al.&nbsp;published the groundbreaking <a href="https://arxiv.org/abs/1312.5602">Playing Atari with Deep Reinforcement Learning</a> in <em>2013</em>, claiming superhuman performance. Surely someone would have noticed by now?</p>
<p>Apparently not, and then most DRL algorithms for the next six years used either the same human scores reported in that paper, or human beginners. It’s true that DQN significantly outperformed their own human player, but that player was not, by far, <em>the best in the world</em>. Other recent claims of superhuman performance have proven that claim against the best players in the world (the paper mentions AlphaGo against Lee Sedol, OpenAI Five against OG, and AlphaStar against Mana), but not for the Atari benchmark.</p>
<p>The most poignant detail to me in this paper involved the common “normalized human score”, where 0% is the score of a random agent, and 100% is the score of the human baseline. <em>On this scale, the median score achieved by the world record holders across all Atari games is 4.4k%</em>. Clearly you can’t claim superhuman performance if there are humans who beat your target by a factor of 44, unless you yourself exceed this score.</p>
<p>For reference, the original Rainbow algorithm achieved a median of 200% over all Atari games, and other algorithms seem to do worse. If the normalized human score is fitted to a maximum equal to the human world record for each game, and run with different time limits, a tuned IQN variant of Rainbow receives a median score of less than 4% (there were other problems with the way benchmarks were done, and correcting for them reduces performance even further).</p>
<p>We have a long way to go then. The paper has a useful analysis drawing on both previous and original research as to <em>why</em> DRL algorithms are so bad at Atari, and I encourage a careful reading. Some of them, such as reward clipping, are called out in previous research as explicitly chosen to improve performance, but (to treat this particular example), it has been mentioned that this causes the agent to prefer many small rewards over a single large reward.</p>
<p>I encourage anyone working with the Atari benchmark to read the paper for themselves.</p>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li>I actually find it somewhat personally encouraging that there’s room for improvement on Atari. It’s easy to experiment, and I have some ideas myself.</li>
<li>That said, it is rather scary that we could overlook something like this for so long, as a community.</li>
<li>Anyway, <em>someone</em> will take this as a call to arms, and make progress. Peter Drucker said, “If you can’t measure it, you can’t improve it.” Now that we have better measurements, I predict improvements.</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/drl-not-superhuman-on-atari/</guid>
  <pubDate>Sun, 18 Aug 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/drl-not-superhuman-on-atari.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>Three Method Comparison for Traffic Signal Control</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/three-method-comparison-for-traffic-signal-control/</link>
  <description><![CDATA[ 





<section id="this-week" class="level1">
<h1>This week</h1>
<p>This week’s paper, <a href="https://arxiv.org/abs/1908.02673v1">Large-scale traffic signal control using machine learning: some traffic flow considerations</a>, caught my eye for several reasons. First, traffic signal control is relevant to my own group’s work involving microservice and network traffic management. Second, the authors use cellular automaton rule 184 as their traffic model, which is actually the first time I’ve seen a cellular automaton used for something serious since <a href="https://www.wolframscience.com/nks/">A New Kind of Science</a>, despite that book’s claim about the likely broad usefulness of simple programs for complex purposes. Lastly, the authors find that supervised learning and random search outperform deep reinforcement learning for high-occupancies of the traffic flow network,</p>
<blockquote class="blockquote">
<p>For occupancies &gt; 75% during training, DRL policies perform very poorly for all traffic conditions, which means that DRL methods cannot learn under highly congested conditions.</p>
</blockquote>
<p>and that they recommend practitioners <em>throw away</em> congested data!</p>
<blockquote class="blockquote">
<p>Our findings imply that it is advisable for current DRL methods in the literature to discard any congested data when training, and that doing this will improve their performance under all traffic conditions.</p>
</blockquote>
<p>I also have to admit that I’ve thought to myself, waiting at empty intersections for a light to turn green, that I could just <em>solve</em> this problem with DRL. If I’m wrong, that would be very interesting and surprising.</p>
</section>
<section id="considerations-in-a-nutshell" class="level1">
<h1>Considerations in a nutshell</h1>
<p>The introduction and background are well summarized in their last paragraph:</p>
<blockquote class="blockquote">
<p>In summary, most recent studies focus on developing effective and robust multi-agent DRL algorithms to achieve coordination among intersections. The number of intersections in those studies are usually limited, thus their results might not apply to large open network. Although the signal control is indeed a continuing problem, it has been always modeled as an episodic process. From the perspective of traffic considerations, expert knowledge has only been incorporated in down-scaling the size of the control problem or designing novel reward functions for DRL algorithm. Few studies have tested their methods given different traffic demands, or shed lights on the learning performance under different traffic conditions, especially the congestion regimes. To fill the gap, our study will treat the large-scale traffic control as a continuing problem and extend classical RL algorithm to fit it. More importantly, noticing the lack of traffic considerations on learning performance, we will train DRL policies under different density levels and explore the results from a traffic flow perspective.</p>
</blockquote>
</section>
<section id="set-up" class="level1">
<h1>Set up</h1>
<section id="traffic" class="level2">
<h2 class="anchored" data-anchor-id="traffic">Traffic</h2>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="http://atlas.wolfram.com/01/01/184/01_01_108_184.gif#right" class="img-fluid figure-img"></p>
<figcaption>CA Rule 184</figcaption>
</figure>
</div>
<p>This is elementary cellular automaton (CA) rule 184. Elementary cellular automata operate on a binary vector, producing a new binary vector in each step that’s a function of the previous one. For each entry in the previous vector, the new value of the corresponding entry in the resulting vector depends on the previous entry and its neighbors to the left and right. There are 256 possible rules with this formulation, and this picture is of the 184th rule set when ordered in the natural way.</p>
<p>Rule 184 can be thought of as a flow of cars along a lane of traffic. Cars move forward (right) by one cell each step only if there is an open space in front of them, otherwise they wait for one to open up. Here’s an example:</p>
<div id="cell-4" class="cell" data-execution_count="1">
<div class="sourceCode cell-code" id="cb1" style="background: #f1f3f5;"><pre class="sourceCode python code-with-copy"><code class="sourceCode python"><span id="cb1-1"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> rule_184(lane):</span>
<span id="cb1-2">    l <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>] <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">+</span> lane <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">+</span> [<span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>] <span class="co" style="color: #5E5E5E;
background-color: null;
font-style: inherit;"># pad</span></span>
<span id="cb1-3">    <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">return</span> [(l[i<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>] <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">and</span> <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">not</span> l[i]) <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">or</span> (l[i] <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">and</span> l[i<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">+</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>])</span>
<span id="cb1-4">            <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> i <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">range</span>(<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>,<span class="bu" style="color: null;
background-color: null;
font-style: inherit;">len</span>(l)<span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">-</span><span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">1</span>)]</span>
<span id="cb1-5"></span>
<span id="cb1-6"><span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">def</span> show(t, lane):</span>
<span id="cb1-7">    <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">print</span>(<span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">f't</span><span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">{</span>t<span class="sc" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">}</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">:</span><span class="ch" style="color: #20794D;
background-color: null;
font-style: inherit;">\t</span><span class="ss" style="color: #20794D;
background-color: null;
font-style: inherit;">'</span>, <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">' '</span>.join([<span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'🚘'</span> <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">if</span> i <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">else</span> <span class="st" style="color: #20794D;
background-color: null;
font-style: inherit;">'_'</span> <span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> i <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> lane]) )</span>
<span id="cb1-8"></span>
<span id="cb1-9">ti <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> [<span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">True</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>, <span class="va" style="color: #111111;
background-color: null;
font-style: inherit;">False</span>]</span>
<span id="cb1-10"></span>
<span id="cb1-11"><span class="cf" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">for</span> i <span class="kw" style="color: #003B4F;
background-color: null;
font-weight: bold;
font-style: inherit;">in</span> <span class="bu" style="color: null;
background-color: null;
font-style: inherit;">range</span>(<span class="dv" style="color: #AD0000;
background-color: null;
font-style: inherit;">7</span>):</span>
<span id="cb1-12">    show(i, ti)</span>
<span id="cb1-13">    ti <span class="op" style="color: #5E5E5E;
background-color: null;
font-style: inherit;">=</span> rule_184(ti)</span></code></pre></div>
<div class="cell-output cell-output-stdout">
<pre><code>t0:  🚘 🚘 🚘 🚘 🚘 _ _ 🚘 _ _ _ _ _ _ _
t1:  🚘 🚘 🚘 🚘 _ 🚘 _ _ 🚘 _ _ _ _ _ _
t2:  🚘 🚘 🚘 _ 🚘 _ 🚘 _ _ 🚘 _ _ _ _ _
t3:  🚘 🚘 _ 🚘 _ 🚘 _ 🚘 _ _ 🚘 _ _ _ _
t4:  🚘 _ 🚘 _ 🚘 _ 🚘 _ 🚘 _ _ 🚘 _ _ _
t5:  _ 🚘 _ 🚘 _ 🚘 _ 🚘 _ 🚘 _ _ 🚘 _ _
t6:  _ _ 🚘 _ 🚘 _ 🚘 _ 🚘 _ 🚘 _ _ 🚘 _</code></pre>
</div>
</div>
<p>The cellular automaton simulates a lane of traffic, and the authors wire two of these lanes up between each adjacent traffic light to create a grid network. The network is laid out on a torus, so there are no boundaries.</p>
<blockquote class="blockquote">
<p>The signalized network corresponds to a homogeneous grid network of bidirectional streets, with one lane per direction of length <img src="https://latex.codecogs.com/png.latex?n%20=%205"> cells between neighboring traffic lights.</p>
</blockquote>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://computable.ai/static/images/signalized_network.png" class="img-fluid figure-img"></p>
<figcaption>Signalized network</figcaption>
</figure>
</div>
<blockquote class="blockquote">
<p>The connecting links to form the torus are shown as dashed directed links; we have omitted the cells on these links to avoid clutter. Each segment has n = 5 cells; an additional cell has been added downstream of each segment to indicate the traffic light color.</p>
</blockquote>
<p>Cars arriving at a green traffic light choose a random “direction” in which to continue. Green lights are on for a minimum of three steps.</p>
</section>
<section id="learning" class="level2">
<h2 class="anchored" data-anchor-id="learning">Learning</h2>
<p>Each traffic signal is managed by an agent, which has two actions it can take at any time step: turn the light red/green for the North-South approaches, or the opposite. The state observable by each agent is an <img src="https://latex.codecogs.com/png.latex?8%5Ctimes%20n"> matrix of bits corresponding to the four incoming and four outgoing CA vectors, and the output is the probability of turning the light red for the North-South approaches. Only one neural net is actually trained, and used by all agents, since there’s no reason for them to be different in this formulation. For the DRL agent, the reward is the <em>incremental</em> average flow per lane (not the average flow per lane), which the authors mention is lower-variance. The authors use a custom infinite-horizon variant of REINFORCE they call REINFORCE-TD.</p>
</section>
</section>
<section id="experiments" class="level1">
<h1>Experiments</h1>
<p>The authors use a maximum-queue-first (LQF) greedy algorithm as their baseline for comparison, which services the lane with the longest queue length at all times.</p>
<section id="random-policies" class="level2">
<h2 class="anchored" data-anchor-id="random-policies">Random policies</h2>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://computable.ai/static/images/traffic_signals_figure4.png" class="img-fluid figure-img"></p>
<figcaption>Figure 4</figcaption>
</figure>
</div>
<p>They begin by randomly reinitializing the parameters of the neural network, and discover that ~15% of random policies are competitive (that is, they can outperform LQF for some traffic densities). They also note a previously undiscovered pattern that “all policies, no matter how bad, are best when the density exceeds approximately 75%.” How odd.</p>
</section>
<section id="supervised-learning-policies" class="level2">
<h2 class="anchored" data-anchor-id="supervised-learning-policies">Supervised learning policies</h2>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://computable.ai/static/images/traffic_signals_figure5.png" class="img-fluid figure-img"></p>
<figcaption>Figure 5</figcaption>
</figure>
</div>
<p>They then train a policy with supervised learning, and surprisingly, with only the two obvious extreme examples, the resulting policy is near-optimal.</p>
</section>
<section id="drl-policies" class="level2">
<h2 class="anchored" data-anchor-id="drl-policies">DRL policies</h2>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://computable.ai/static/images/traffic_signals_figure6.png" class="img-fluid figure-img"></p>
<figcaption>Figure 6</figcaption>
</figure>
</div>
<blockquote class="blockquote">
<p>Policies trained with constant demand and random initial parameters <img src="https://latex.codecogs.com/png.latex?%5Ctheta">. The label in each diagram gives the iteration number and the constant density value. First column: NS red probabilities of the extreme states, <img src="https://latex.codecogs.com/png.latex?%5Cpi(s1)"> in dashed line and <img src="https://latex.codecogs.com/png.latex?%5Cpi(s2)"> in solid line. The remaining columns show the flow-density diagrams obtained at different iterations, and the last column shows the iteration producing the highest flow at <img src="https://latex.codecogs.com/png.latex?k%20=%200.5">, if not reported on a earlier column.</p>
</blockquote>
<p>Finally, they run two experiments with DRL policies, as described above. These policies seem to do rather poorly in general compared to random search and supervised learning, and as density increases, they stop learning much of anything.</p>
<blockquote class="blockquote">
<p>We conjecture that this result is a consequence of a property of congested urban networks and has nothing to do with the algorithm to train the DRL policy.</p>
</blockquote>
<p>I’m skeptical. See my parting thoughts.</p>
<p>The other experiments the authors perform just confirms that average flow per lane does worse than incremental average flow per lane.</p>
</section>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li>In the end, I’m way more interested in the experimental setup of this paper than the conclusions. As usual, I learned a ton, and I may actually use rule 184 as a model for traffic flow on something.</li>
<li>Isn’t it <em>obvious</em> given their problem formulation that the agents can’t learn under conditions of congestion, since it means their input is essentially whited out? I would be more impressed with the conclusion if a neural net with complete visibility had trouble learning with congestion. It also seems to me <em>extremely</em> suggestive that a supervised policy can learn from only two examples, and I would very much like to see if the major conclusions of this paper explode with a more realistic network topology. Queueing theory contains all sorts of counterintuitive surprises, and it seems likely to me that their results are more indicative of one of those surprises, rather than some deep fact about DRL’s ability to manage urban congestion.</li>
<li>It’s interesting that they formulate the problem as a continuing one, against the prevailing trend in the traffic signal control literature. I agree with them, that even if you get to a state where there’s no traffic, that’s a function of the demand, not of the agent’s choices. I bring this up because I too have found that it’s <em>really quite important</em> to recognize an infinite-horizon problem when you have one, or else your agent learns to rack up debts until the end of the artificial episode when all is “forgiven”.</li>
<li>It’s fascinating that all random policies, no matter how bad, are best around 75% congestion. I have been admonished to avoid scheduling myself at more than 70% capacity to avoid the ringing effect. I wonder if this is an empirical vindication of that…</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/three-method-comparison-for-traffic-signal-control/</guid>
  <pubDate>Sun, 11 Aug 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/three-method-comparison-for-traffic-signal-control.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>Learning Compound and Composable Policies</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/learning-compound-and-composable-policies/</link>
  <description><![CDATA[ 





<section id="this-week" class="level1">
<h1>This week</h1>
<p>Just a sketch this week, of <a href="https://arxiv.org/abs/1905.09668">Hierarchical Reinforcement Learning for Concurrent Discovery of Compound and Composable Policies</a>.</p>
<p>I’ve been hearing hierarchical RL mentioned frequently lately, and while I understand it’s a way to encode human expertise to achieve otherwise intractible goals, it has also seemed a bit like cheating. However, I have a day job, and this serves as a healthy dose of pragmatism. I also think that even when the goal is fundamental progress, it’s often a good idea to achieve the goal <em>in any way possible</em>, and then follow-up by working the cheats out of the system one by one. So when I read the abstract of this paper, I was feeling more receptive than previously.</p>
<p>Part of what made hierarchical RL seem not worth the cheating was how kludgy and inefficient the usual methods were, retraining a whole new policy from scratch for each subtask. That’s why this week’s paper caught my eye:</p>
<blockquote class="blockquote">
<p>… we propose an algorithm for learning both compound and composable policies <strong>within the same learning process</strong> by exploiting the off-policy data generated from the compound policy.</p>
</blockquote>
<p>Their resulting algorithm, “Hierarchical Intentional-Unintentional Soft Actor-Critic” (HIU-SAC), efficiently trains all sub-policies simultaneously, choosing actions to perform in the environment using a weighted average of the “votes” of all sub-policies, with weights given by a learned selector network (which is <em>also</em> simultaneously trained).</p>
</section>
<section id="composable-hierarchical-rl" class="level1">
<h1>Composable hierarchical RL</h1>
<section id="architecture" class="level2">
<h2 class="anchored" data-anchor-id="architecture">Architecture</h2>
<p><img alt="Hierarchical policy diagram" src="https://computable.ai/static/images/policy_network.png#right" height="300px" width="300px" style="margin: 10px"></p>
<p>The composite policy consists of the individual policy networks, each with its own reward function, trained to take observations <img src="https://latex.codecogs.com/png.latex?s"> in and output parameters of a conditional Gaussian. There is also a special activation vector selector network trained on the same states to produce weights corresponding to how much each constituent policy applies to the current state. All of these networks share early layers, since they all benefit from an accurate high-level state representation. Finally, some function <img src="https://latex.codecogs.com/png.latex?f"> takes all of these outputs and determines what action <img src="https://latex.codecogs.com/png.latex?a"> to <em>actually</em> take in the environment.</p>
<p><img alt="Q-value function diagram" src="https://computable.ai/static/images/q_fcn_network.png#left" height="250px" width="250px" style="margin: 10px"></p>
The Q function networks are similarly arranged, sharing early layers which take a state <img src="https://latex.codecogs.com/png.latex?s"> and an action <img src="https://latex.codecogs.com/png.latex?a"> to produce a Q function for each subtask, as well as a composite Q function.
<div style="clear:both">
&nbsp;
</div>
</section>
<section id="simultaneous-learning" class="level2">
<h2 class="anchored" data-anchor-id="simultaneous-learning">Simultaneous learning</h2>
<blockquote class="blockquote">
<p>Most methods learn the composable tasks one at a time, and later, the compound task. This procedure is not scalable as all the experience collected for each learning process is only used for that specific process. Also, it is not possible to start learning more complex tasks unless all the compos- able policies have been successfully learned. The method proposed in this section is based on the idea that a single stream of experience can be used to improve not only the policy that is generating the behavior but also, indirectly, many other policies.</p>
</blockquote>
<p>The authors refer to the composite policy acting as the “intentional” policy (the “behavior” policy in an off-policy setting), and the composable sub-policies as the “unintentional” policies (each one a “target” policy in an off-policy setting). They use a variation on SAC to train the composite and composable policies simultaneously within the maximum entropy framework.</p>
<p>The objective function for the Q networks simply maximize the expected sum of all mean-squared Bellman errors for each Q network, for each tuple in the replay buffer <img src="https://latex.codecogs.com/png.latex?%5Cmathcal%7BD%7D">. The objective function for the policy is simply the sum of the objective functions for each intentional and unintentional policy. Each policy objective optimizes the expected difference for each state in <img src="https://latex.codecogs.com/png.latex?%5Cmathcal%7BD%7D"> between the Q value and log-probability of the selected action (adjustable by temperature <img src="https://latex.codecogs.com/png.latex?%5Calpha">), over all possible actions. HIU-SAC then alternates between policy evaluation and policy improvement steps following SAC.</p>
</section>
<section id="the-importance-of-maximizing-entropy-to-adequate-exploration" class="level2">
<h2 class="anchored" data-anchor-id="the-importance-of-maximizing-entropy-to-adequate-exploration">The importance of maximizing entropy to adequate exploration</h2>
<p>It is interesting that the entropy-maximizing RL objective was <em>absolutely necessary</em> for exploring broadly enough to train all of these policies at once.</p>
<blockquote class="blockquote">
<p>Note that populating the replay memory buffer with rich experiences is essential for acquiring multiple skills in an off-policy manner. The composable policies learned unintentionally had similar performance than the policies obtained in single-task formulations only when the compound policy was able to efficiently explore the environment. For this reason, the algorithm was built on a maximum entropy RL framework to favor exploration during the learning process.</p>
</blockquote>
</section>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li>In a way, the methods proposed here seem rather obvious, and I found this paper quite easy to understand given that it violated none of my expectations. I also haven’t been paying enough attention to hierarchical RL to know off-hand why training the sub-policies in parallel off of the same recorded environment interactions hasn’t been tried before (or whether it has been without my notice). Perhaps it was necessary for off-policy RL to reach a level of maturity sufficient for sub-policies to see enough relevant data to train? In any case, don’t hear me faulting the authors for trying the obvious. It is relieving a <em>non</em>-obvious that a straightforward formulation works so well.</li>
<li>I’d love to see this work combined with imitation learning and inverse RL to figure out what sub-policies are necessary in the first place from demonstrations. That seems like a very practical framework for real-world learning.</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/learning-compound-and-composable-policies/</guid>
  <pubDate>Sun, 04 Aug 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/learning-compound-and-composable-policies.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>Efficient exploration with self-imitation learning</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/efficient-exploration-with-self-imitation-learning/</link>
  <description><![CDATA[ 





<section id="this-week" class="level1">
<h1>This week</h1>
<p>Several paper caught my eye this week, but I’ll be discussing only <a href="https://arxiv.org/abs/1907.10247">Efficient Exploration with Self-Imitation Learning via Trajectory-Conditioned Policy</a> in more depth. I’m choosing this paper because, as happens sometimes, I had this idea myself a few weeks ago. It’s especially exciting to see something you suspected might improve the world fleshed out and vindicated.</p>
<p>This is the basic form of my shower-throught idea:</p>
<blockquote class="blockquote">
<p>This paper investigates the imitation of diverse past trajectories and how that leads [to] further exploration and avoids getting stuck at a sub-optimal behavior. Specifically, we propose to use a buffer of the past trajectories to cover diverse possible directions. Then we learn a trajectory-conditioned policy to imitate any trajectory from the buffer, treating it as a demonstration. After completing the demonstration, the agent performs random exploration.</p>
</blockquote>
</section>
<section id="the-problem" class="level1">
<h1>The problem</h1>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://computable.ai/static/images/maze_icon_map.png#right" class="img-fluid figure-img"></p>
<figcaption>Maze</figcaption>
</figure>
</div>
<p>The main problem the authors want to solve is insufficient exploration leading to a sub-optimal policy. If you don’t explore your environment enough, you will find local rewards, but miss globally optimal rewards. In this maze (their Figure 1), you can see that an agent that fails to explore will collect two apples in the next room, but may miss acquiring the key, unlocking the door, collecting an apple, and discovering the treasure.</p>
<p>In the notoriously difficult Atari game (for RL agents) Montezuma’s Revenge, it is similarly extremely unlikely that random exploration suffices to explore the environment and achieve a high score. The authors report state-of-the-art performance without expert demonstrations on Montezuma’s Revenge, netting 25k points.</p>
</section>
<section id="sota-without-demonstrations" class="level1">
<h1>SOTA without demonstrations</h1>
<p>So, more precisely, how did they achieve this, and why does it work?</p>
<blockquote class="blockquote">
<p>The main idea of our method is to maintain a buffer of diverse trajectories collected during training and to train a trajectory-conditioned policy by leveraging reinforcement learning and supervised learning to roughly follow demonstration trajectories sampled from the trajectory buffer. Therefore, the agent is encouraged to explore beyond various visited states in the environment and gradually push its exploration frontier further… We name our method as Diverse Trajectory-conditioned Self-Imitation Learning (DTSIL).</p>
</blockquote>
<section id="the-trajectory-buffer" class="level2">
<h2 class="anchored" data-anchor-id="the-trajectory-buffer">The trajectory buffer</h2>
<p>Their trajectory buffer <img src="https://latex.codecogs.com/png.latex?%5Cmathcal%7BD%7D"> contains <img src="https://latex.codecogs.com/png.latex?N"> 3-tuples <img src="https://latex.codecogs.com/png.latex?%5C%7B%5Cleft(e%5E%7B(1)%7D,%20%5Ctau%5E%7B(1)%7D,%20n%5E%7B(1)%7D%5Cright),%20%5Cleft(e%5E%7B(2)%7D,%20%5Ctau%5E%7B(2)%7D,%20n%5E%7B(2)%7D%5Cright),%20%5Cldots%20%5Cleft(e%5E%7B(N)%7D,%20%5Ctau%5E%7B(N)%7D,%20n%5E%7B(N)%7D%5Cright)%20%5C%7D"> where <img src="https://latex.codecogs.com/png.latex?e%5E%7B(i)%7D"> is a high-level state representation, <img src="https://latex.codecogs.com/png.latex?%5Ctau%5E%7B(i)%7D"> is the shortest trajectory achieving the highest reward and arriving at <img src="https://latex.codecogs.com/png.latex?e%5E%7B(i)%7D">, and <img src="https://latex.codecogs.com/png.latex?n%5E%7B(i)%7D"> is the number of times <img src="https://latex.codecogs.com/png.latex?e%5E%7B(i)%7D"> has been encountered. Whenever they roll out a new episode, they check each high-level state representation encountered against those in <img src="https://latex.codecogs.com/png.latex?%5Cmathcal%7BD%7D">, increment <img src="https://latex.codecogs.com/png.latex?n">, and if <img src="https://latex.codecogs.com/png.latex?%5Ctau"> is better they replace <img src="https://latex.codecogs.com/png.latex?%5Ctau"> for that entry.</p>
</section>
<section id="sampling" class="level2">
<h2 class="anchored" data-anchor-id="sampling">Sampling</h2>
<p>When training their trajectory-conditioned policy, they sample each 3-tuple with weight <img src="https://latex.codecogs.com/png.latex?%7B1%7D%5Cover%7B%5Csqrt%7Bn%5E%7B(i)%7D%7D%7D">. Notice that this will cause them to sample <em>less</em> frequently-visited states more often, encouraging exploration.</p>
</section>
<section id="imitation-reward" class="level2">
<h2 class="anchored" data-anchor-id="imitation-reward">Imitation reward</h2>
<p>Given a trajectory <img src="https://latex.codecogs.com/png.latex?g"> sampled from the buffer, and during interaction with the environment, the agent receives a positive reward if the current state has an embedding within some <img src="https://latex.codecogs.com/png.latex?%5CDelta%20t"> of the current timestep in <img src="https://latex.codecogs.com/png.latex?g">. Otherwise the imitation reward is 0. Once it reaches the end of <img src="https://latex.codecogs.com/png.latex?g">, there is no further imitation reward, and it explores randomly. The imitation reward is one of two components of the <img src="https://latex.codecogs.com/png.latex?r%5E%7BDTSIL%7D_%7Bt%7D"> RL reward, where the other is a simple monotonic function of the reward received at each timestep.</p>
</section>
<section id="policy-architecture" class="level2">
<h2 class="anchored" data-anchor-id="policy-architecture">Policy architecture</h2>
<p>The DTSIL policy architecture is recurrent and attentional, inspired by machine translation!</p>
<blockquote class="blockquote">
<p>Inspired by neural machine translation methods, the demonstration trajectory is the source sequence and the incomplete trajectory of the agent’s state representations is the target sequence. We apply a recurrent neural network and an attention mechanism to the sequence data to predict actions that would make the agent to follow the demonstration trajectory.</p>
</blockquote>
</section>
<section id="rl-objective" class="level2">
<h2 class="anchored" data-anchor-id="rl-objective">RL objective</h2>
<p>DTSIL is trained using a policy gradient algorithm (PPO, in their experiments), and RL loss</p>
<p><img src="https://latex.codecogs.com/png.latex?%5Cmathcal%20L%5E%7BRL%7D%20=%20%7B%5Cmathbb%7BE%7D%7D_%7B%5Cpi_%5Ctheta%7D%20%5B-%5Clog%20%5Cpi_%5Ctheta(a_t%7Ce_%7B%5Cleq%20t%7D,%20o_t,%20g)%20%5Cwidehat%7BA%7D_t%5D"></p>
<p>where <img src="https://latex.codecogs.com/png.latex?%5Cwidehat%7BA%7D_t=%5Csum%5E%7Bn-1%7D_%7Bd=0%7D%20%5Cgamma%5E%7Bd%7Dr%5E%5Ctext%7BDTSIL%7D_%7Bt+d%7D%20+%20%5Cgamma%5En%20V_%5Ctheta(e_%7B%5Cleq%20t+n%7D,%20o_%7Bt+n%7D,%20g)%20-%20V_%5Ctheta(e_%7B%5Cleq%20t%7D,%20o_t,%20g)"></p>
</section>
<section id="sl-objective" class="level2">
<h2 class="anchored" data-anchor-id="sl-objective">SL objective</h2>
<p>In each parameter optimization step, they also include a supervised loss designed to maximize the log probability of taking an action that imitates the chosed demonstration exactly to better leverage a past trajectory <img src="https://latex.codecogs.com/png.latex?g">.</p>
<p><img src="https://latex.codecogs.com/png.latex?%5Cmathcal%20L%5E%5Ctext%7BSL%7D%20=%20-%20%5Clog%20%5Cpi_%5Ctheta(a_t%7Ce_%7B%5Cleq%20t%7D,%20o_t,%20g)%20%5Ctext%7B,%20where%20%7D%20g%20=%20%5C%7Be_0,%20e_1,%20%5Ccdots,%20e_%7B%7Cg%7C%7D%5C%7D"></p>
</section>
<section id="optimization" class="level2">
<h2 class="anchored" data-anchor-id="optimization">Optimization</h2>
<p>The final parameter update is thus</p>
<p><img src="https://latex.codecogs.com/png.latex?%5Ctheta%20%5Cgets%20%5Ctheta%20-%20%5Ceta%20%5Cnabla_%5Ctheta%20(%5Cmathcal%7BL%7D%5E%5Ctext%7BRL%7D+%5Cbeta%20%5Cmathcal%7BL%7D%5E%5Ctext%7BSL%7D)"></p>
</section>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li>I <em>love</em> seeing methods developed for generative language models used in another context entirely, to generate another kind of sequence. I’m overjoyed that it worked well.</li>
<li>They need a high-level embedding for two reasons: first because storing entire trajectories exactly in memory is expensive, and second because it’s quite difficult to re-execute a previously-encountered trajectory exectly, so in order for this method to work at all it’s important that an <em>approximate</em> re-execution be possible.</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/efficient-exploration-with-self-imitation-learning/</guid>
  <pubDate>Sun, 28 Jul 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/efficient-exploration-with-self-imitation-learning.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>Keeping to the Narrow Path</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/keeping-to-the-narrow-path/</link>
  <description><![CDATA[ 





<section id="this-week" class="level1">
<h1>This week</h1>
<p>This week’s highlight is a paper on imitation learning: <a href="https://arxiv.org/abs/1907.05634">Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative Sampling</a>, chosen again for pragmatic reasons. The problem my team is currently working on has both reasons for wanting high sample efficiency: training would be prohibitively slow without something to kickstart it, and actions taken in the real world can get expensive.</p>
<p>I know I said I’d be experimenting with shorter, more bite-sized posts, but… next time. (If you want that, you can just stop reading after the “Key intuition” section.)</p>
</section>
<section id="the-problem" class="level1">
<h1>The problem</h1>
<p>Learning from demonstrations is more difficult than it may seem at first glance. The trouble mainly stems from covariate shift: the input distribution your agent will see in production is very likely to be different than that encountered during training. Many machine learning algorithms have this problem, reinforcement learning algorithms included, but imitation learning has it especially bad, for a simple reason: the expert demonstrations you are attempting to follow necessarily explore a very small subset of the state space. The whole <em>point</em> of them is to stay on good trajectories, meaning bad trajectories never get explored.</p>
<p>This causes two issues:</p>
<ol type="1">
<li>The agent can’t in general figure out how to get back into the subset of state space where the expert demonstrations apply, even if it gets only slightly off-course, and</li>
<li>Value functions for states and actions are affected by unseen states, making it very <em>likely</em> that the agent will wander off as soon as it’s allowed.</li>
</ol>
</section>
<section id="key-intuition" class="level1">
<h1>Key intuition</h1>
<p>The authors solve this problem by pre-training with supervised learning using a loss function that drives down the value of all states outside of those explored in the expert demonstrations <img src="https://latex.codecogs.com/png.latex?U">, by an amount proportional to their Euclidean distance from the closest state in <img src="https://latex.codecogs.com/png.latex?U">. In their own words:</p>
<blockquote class="blockquote">
<p>Consider a state <img src="https://latex.codecogs.com/png.latex?s"> in the demonstration and its nearby state <img src="https://latex.codecogs.com/png.latex?%5Ctilde%7Bs%7D"> that is not in the demonstration. The key intuition is that <img src="https://latex.codecogs.com/png.latex?%5Ctilde%7Bs%7D"> should have a lower value than <img src="https://latex.codecogs.com/png.latex?s">, because otherwise <img src="https://latex.codecogs.com/png.latex?%5Ctilde%7Bs%7D"> likely should have been visited by the demonstrations in the first place. If a value function has this property for most of the pair <img src="https://latex.codecogs.com/png.latex?(s,%5Ctilde%7Bs%7D)"> of this type, the corresponding policy will tend to correct its errors by driving back to the demonstration states because the demonstration states have locally higher values.</p>
</blockquote>
<p>And Figure 1 is a nice visual demonstration:</p>
<p><a href="../../static/images/VINS_Figure_1.jpeg"><img alt="VINS Figure 1" src="https://computable.ai/static/images/VINS_Figure_1.jpeg"></a></p>
</section>
<section id="value-iteration-with-negative-sampling-vins" class="level1">
<h1>Value Iteration with Negative Sampling (VINS)</h1>
<p>Into the weeds now.</p>
<section id="self-correctable-policy" class="level2">
<h2 class="anchored" data-anchor-id="self-correctable-policy">Self-correctable policy</h2>
<p>The first bit of their algorithm is the definition of their self-correcting policy. It’s essentially a formalization of what we said above about <img src="https://latex.codecogs.com/png.latex?s"> and <img src="https://latex.codecogs.com/png.latex?%5Ctilde%7Bs%7D">.</p>
<p>If <img src="https://latex.codecogs.com/png.latex?s%20%5Cin%20U"> (if <img src="https://latex.codecogs.com/png.latex?s"> is in the expert demonstrations), then <img src="https://latex.codecogs.com/png.latex?V(s)%20=%20V%5E%7B%5Cpi_e%7D(s)%20%5Cpm%20%5Cdelta_V"> (“just what the value would be in the expert demonstrations, plus some error”).</p>
<p>But if <img src="https://latex.codecogs.com/png.latex?s%20%5Cnot%5Cin%20U">, <img src="https://latex.codecogs.com/png.latex?V(s)%20=%20V%5E%7B%5Cpi_e%7D(%5CPi_U(s))%20-%20%5Clambda%20%5C%7Cs-%5CPi_U(s)%5C%7C%20%5Cpm%20%5Cdelta_V"> (where <img src="https://latex.codecogs.com/png.latex?%5CPi_U"> gives the closest <img src="https://latex.codecogs.com/png.latex?s%20%5Cin%20U">, so <img src="https://latex.codecogs.com/png.latex?V(s)"> is “the value of the closest <img src="https://latex.codecogs.com/png.latex?s%20%5Cin%20U">, <em>minus the distance to that</em> <img src="https://latex.codecogs.com/png.latex?s%20%5Cin%20U">, plus some error”)</p>
<p>Then the induced policy from this value function is <img src="https://latex.codecogs.com/png.latex?%5Cpi(s)%20%5Ctriangleq%20%5Cunderset%7Ba:%20%5C%7Ca-%5Cpi_%7BBC%7D(s)%5C%7C%5Cle%20%5Czeta%7D%7B%5Coperatorname%7Bargmax%7D%7D%20~V(M(s,%20a))"></p>
<p>Where <img src="https://latex.codecogs.com/png.latex?M(s,a)"> is a learned dynamical model of the environment that gives the next state given the current state and action. <img src="https://latex.codecogs.com/png.latex?%5Cpi_%7BBC%7D(s)"> is the “behavioral clone” policy from the expert demonstrations.</p>
</section>
<section id="rl-algorithm" class="level2">
<h2 class="anchored" data-anchor-id="rl-algorithm">RL algorithm</h2>
<p>To actually achieve <img src="https://latex.codecogs.com/png.latex?V(M(s,a))"> with the necessary properties, they select a state <img src="https://latex.codecogs.com/png.latex?s"> from the demonstrations, perturb it a bit to get <img src="https://latex.codecogs.com/png.latex?%5Ctilde%7Bs%7D"> nearby, and use the original state <img src="https://latex.codecogs.com/png.latex?s"> to approximate <img src="https://latex.codecogs.com/png.latex?%5CPi_U(%5Ctilde%7Bs%7D)"> in the following loss function.</p>
<p><img src="https://latex.codecogs.com/png.latex?%5Cmathcal%7BL%7D_%7Bns%7D(%5Cphi)=%20%5Cmathbf%7BE%7D_%7Bs%20%5Csim%20%5Crho%5E%7B%5Cpi_e%7D,%20%5Ctilde%7Bs%7D%20%5Csim%20perturb(s)%7D%20%5Cleft(V_%7B%5Cbar%20%5Cphi%7D(s)%20-%20%5Clambda%20%5C%7Cs-%5Ctilde%7Bs%7D%5C%7C-%20V_%5Cphi(%5Ctilde%7Bs%7D)%20%5Cright)%5E2"></p>
<p>Finally, here’s the algorithm that uses this and the earlier policy definition:</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://computable.ai/static/images/VINS_Algorithm_2.jpeg#center" class="img-fluid figure-img"></p>
<figcaption>VINS Algorithm 2</figcaption>
</figure>
</div>
</section>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li>I thought it was quite strange that they learned <img src="https://latex.codecogs.com/png.latex?V(s)"> and a dynamical model <img src="https://latex.codecogs.com/png.latex?M(s,a)">, and then used <img src="https://latex.codecogs.com/png.latex?V(M(s,a))"> in the algorithm. I thought, “Why not just learn <img src="https://latex.codecogs.com/png.latex?Q">?” The answer was given in their Section A appendix, and was quite interesting. I’m not sure it applies to our case, but it’s important. TL;DR <img src="https://latex.codecogs.com/png.latex?Q(s,a)"> learned from demonstrations <em>alone</em> is degenerate, because there’s always a <img src="https://latex.codecogs.com/png.latex?Q"> that perfectly matches the demonstrations <em>and doesn’t depend at all on</em> <img src="https://latex.codecogs.com/png.latex?a">.</li>
<li>One of my coworkers (and upcoming Computable author!) wondered to me if the induced policy could be made explicit, by explicitly training a policy network to bring the agent back into safe territory. It could be trained with gradient descent, because <img src="https://latex.codecogs.com/png.latex?V(M(s,a))"> are just networks, and the technique for training deterministic policies just follows the gradient of the <img src="https://latex.codecogs.com/png.latex?Q"> function. I wonder too.</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/keeping-to-the-narrow-path/</guid>
  <pubDate>Sun, 21 Jul 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/keeping-to-the-narrow-path.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>Look at This: Where We See Shapes, AI Sees Textures</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/look-at-this-where-we-see-shapes-ai-sees-textures/</link>
  <description><![CDATA[ 





<section id="new-series" class="level1">
<h1>New Series</h1>
<p><img src="http://weknowmemes.com/wp-content/uploads/2011/12/look-at-this-duck.jpg#right" style="margin-left:15px" width="350" height="268"></p>
<p>We’re starting a simple new series called Look at This, where we briefly plug an article that taught us something.</p>
<p>Our first highlight will be a Quanta article about what CNNs learn when trained in “the usual way”:</p>
<p><a href="https://www.quantamagazine.org/where-we-see-shapes-ai-sees-textures-20190701/">Where We See Shapes, AI Sees Textures</a></p>
</section>
<section id="textures-not-shapestraining-a-cnn-for-object-recognition-typically-involves-only-showing-the-algorithm-many-examples-of-images-that-contain-or-dont-contain-a-target-object.-humans-also-need-to-see-many-examples-of-various-objects-to-get-the-basic-idea.-humans-however-seem-to-have-a-bias-towards-recognition-by-shape-which-is-missing-from-cnns-in-general.-geirhos-bethge-and-their-colleagues-created-images-that-included-two-conflicting-cues-with-a-shape-taken-from-one-object-and-a-texture-from-another-the-silhouette-of-a-cat-colored-in-with-the-cracked-gray-texture-of-elephant-skin-for-instance-or-a-bear-made-up-of-aluminum-cans-or-the-outline-of-an-airplane-filled-with-overlapping-clock-faces.-presented-with-hundreds-of-these-images-humans-labeled-them-based-on-their-shape-cat-bear-airplane-almost-every-time-as-expected.-four-different-classification-algorithms-however-leaned-the-other-way-spitting-out-labels-that-reflected-the-textures-of-the-objects-elephant-can-clock.this-is-a-problem-worth-solving-since-the-addition-of-even-a-small-amount-of-noise-can-throw-off-cnn-based-classifiers-where-humans-arent-fooled.-adversarial-examples-even-do-this-maliciously-adding-exactly-the-right-amount-of-noise-to-cause-misclassification.-so-how-to-fix-this-geirhos-wanted-to-see-what-would-happen-when-the-team-forced-their-models-to-ignore-texture.-the-team-took-images-traditionally-used-to-train-classification-algorithms-and-painted-them-in-different-styles-essentially-stripping-them-of-useful-texture-information.-when-they-retrained-each-of-the-deep-learning-models-on-the-new-images-the-systems-began-relying-on-larger-more-global-patterns-and-exhibited-a-shape-bias-much-more-like-that-of-humans.the-illustration-of-images-painted-with-alien-textures-that-originally-appeared-here-is-no-longer-hosted-by-quanta-see-the-original-article-for-the-imagery.there-were-many-other-insights-in-this-relatively-short-article-and-i-commend-it-to-you.-it-enriched-my-understanding-of-whats-going-on-in-neural-networks-and-how-far-we-still-need-to-go-to-reach-parity-with-humans." class="level1">
<h1>Textures, not shapesTraining a CNN for object recognition typically involves only showing the algorithm many examples of images that contain or don’t contain a target object. Humans also need to see many examples of various objects to get the basic idea. Humans, however, seem to have a bias towards recognition by <em>shape</em> which is missing from CNNs in general.&gt; Geirhos, Bethge and their colleagues created images that included two conflicting cues, with a shape taken from one object and a texture from another: the silhouette of a cat colored in with the cracked gray texture of elephant skin, for instance, or a bear made up of aluminum cans, or the outline of an airplane filled with overlapping clock faces. Presented with hundreds of these images, humans labeled them based on their shape — cat, bear, airplane — almost every time, as expected. Four different classification algorithms, however, leaned the other way, spitting out labels that reflected the textures of the objects: elephant, can, clock.This is a problem worth solving, since the addition of even a small amount of noise can throw off CNN-based classifiers, where humans aren’t fooled. “Adversarial examples” even do this maliciously, adding exactly the right amount of noise to cause misclassification. So how to fix this?&gt; Geirhos wanted to see what would happen when the team forced their models to ignore texture. The team took images traditionally used to train classification algorithms and “painted” them in different styles, essentially stripping them of useful texture information. When they retrained each of the deep learning models on the new images, the systems began relying on larger, more global patterns and exhibited a shape bias much more like that of humans.<em>(The illustration of images painted with alien textures that originally appeared here is no longer hosted by Quanta — see <a href="https://www.quantamagazine.org/where-we-see-shapes-ai-sees-textures-20190701/">the original article</a> for the imagery.)</em>There were many other insights in this relatively short article, and I commend it to you. It enriched my understanding of what’s going on in neural networks, and how far we still need to go to reach parity with humans.</h1>


</section>

 ]]></description>
  <category>Look at This</category>
  <guid>https://computable.ai/posts/look-at-this-where-we-see-shapes-ai-sees-textures/</guid>
  <pubDate>Tue, 16 Jul 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/look-at-this-where-we-see-shapes-ai-sees-textures.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>Way Off-Policy Batch DRL</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/way-off-policy-batch-drl/</link>
  <description><![CDATA[ 





<section id="this-week" class="level1">
<h1>This week</h1>
<p>Only one paper this week, <em>not</em> because <a href="https://arxiv.org/abs/1905.04819">others</a> failed to catch my eye, but for brevity. Let me know in the comments if you agree that shorter or more focused articles are more attractive. So this week I’ll be examining just one paper: <a href="https://arxiv.org/abs/1907.00456">Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog</a>. As with last week’s papers, this week’s is interesting to me professionally. Batch DRL is a way to solve the sample efficiency problem, from a certain perspective. It’s mostly the online learning that costs too much when sample efficiency is low, so solving the problems that come with attempting to train offline might allow us to do many of the same things we could do if we had high online sample efficiency.</p>
</section>
<section id="rl-for-open-domain-dialog-generation" class="level1">
<h1>RL for open-domain dialog generation</h1>
<p>The author’s domain is dialog generation. They want to build a better chat bot, and they have quite a few recorded conversations. RL is good at refining these processes, but has a cold-start problem, plus they would certainly prefer to make use of the data they have on-hand. For this, they need to be able to make use of offline data, hence “<em>Way</em> Off-Policy”. This data is so off-policy it wasn’t even <em>generated</em> by a policy.</p>
<p>So they want to train DRL from samples acquired from some other control of the system (in their case, human interaction data), much like <a href="https://arxiv.org/abs/1704.03732">Deep Q-learning from Demonstrations</a>. There are a couple of reasons this is important for others such as myself:</p>
<blockquote class="blockquote">
<p>First, since collecting real-world interaction data can be expensive and time-consuming, algorithms must be able to leverage off-policy data - collected from vastly different systems, far into the past - in order to learn.</p>
</blockquote>
<blockquote class="blockquote">
<p>Second, it is often necessary to carefully test a policy before deploying it to the real world; for example, to ensure its behavior is safe and appropriate for humans. Thus the algorithm must be able to learn offline first, from a static batch of data, without the ability to explore</p>
</blockquote>
</section>
<section id="a-generative-model-q-learning" class="level1">
<h1>A generative model + Q learning</h1>
<p>The authors first pre-train a generative model on the distribution of collected trajectories, and initialize the Q networks from this model. They then sample a fixed number of actions from it, and output the one with the highest Q-value as their policy’s decision. In later reinforcement learning, they penalize their model for KL-divergence from this distribution.</p>
<blockquote class="blockquote">
<p>To perform batch Q-learning, we first pre-train a generative model of <img src="https://latex.codecogs.com/png.latex?p(a%7Cs)"> using a set of known environment trajectories. In our case, this model is then used to generate the batch data via human interaction. The weights of the Q-network and target Q-network are initialized from the pre-trained model, which helps reduce variance in the Q-estimates and works to combat overestimation bias. To train <img src="https://latex.codecogs.com/png.latex?Q_%7B%CE%B8_%CF%80%7D"> we sample &lt; <img src="https://latex.codecogs.com/png.latex?s_t">, <img src="https://latex.codecogs.com/png.latex?a_t">, <img src="https://latex.codecogs.com/png.latex?r_t">, <img src="https://latex.codecogs.com/png.latex?s_%7Bt+1%7D"> &gt; tuples from the batch, and update the weights of the Q-network to approximate Eq. 1. This forms our baseline model, which we call Batch Q</p>
</blockquote>
</section>
<section id="overestimation-bias" class="level1">
<h1>Overestimation bias</h1>
<blockquote class="blockquote">
<p>Most deep RL algorithms fail to learn from data that is not heavily correlated with the current policy. Even models based on off-policy algorithms lik Q-learning fail to learn when the model is not able to explore during training. This is due to the fact that such algorithms are inherently optimistic in the face of uncertainty.</p>
</blockquote>
<p>If you’re taking the <code>max</code> of something (as in Bellman-equation-based algorithms), then the higher the variance, the higher the <code>max</code> value. This causes an over-estimation bias. We may have seen a really high value for some state once, so now we over-value that state, despite it being atypical. It may not be immediately obvious why this is a <em>problem</em>, but which states are we likely to overvalue? Precisely the states we haven’t visited often. Why is <em>that</em> a problem? This sounds good for exploration, right? But if we’re trying to train our agent with canned data, it’s important that the live agent stick pretty close to the states where the canned data does well, and it’s counter-productive to have it believe that everywhere <em>but</em> the pre-explored state space is worth exploring.</p>
<p>A popular solution to the overestimation problem in Q-learning algorithms is to train <em>two</em> Q networks on the same data, put the input through both, and take the minimum value. This helps with the bias because they’ll likely disagree unless we can be really <em>certain</em> of the value of the input, and if they disagree we can go with the least confident. The authors of the current paper take a different tack, training a single neural net with dropout, and using the disagreement with different dropout masks as an estimate of uncertainty.</p>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li><p>I didn’t talk much about their model architecture, which is “Variational Hierarchical Recurrent Encoder Decoder (VHRED)”, largely because I think if I ever tried to make use of this directly I would employ transformers instead. They do mention that transformer architectures are a “powerful alternative”, but they chose to work with hierarchical architectures so they could extend their work to hierarchical control in the future. That’s interesting. In my own work at the moment, the important thing is the “way off-policy” part, not so much the chat bot part.</p></li>
<li><p>It’s very interesting to me that both of the methods for correcting overestimation bias make use of uncertainty estimators that I’ve seen mentioned elsewhere:</p></li>
</ol>
<ul>
<li><a href="https://arxiv.org/abs/1905.09638">Estimating Risk and Uncertainty in Deep Reinforcement Learning</a></li>
</ul>
<blockquote class="blockquote">
<p>…we show that the disagreement between only two neural networks is sufficient to produce a low-variance estimate of the epistemic uncertainty on the return distribution, thus providing a simple and computationally cheap uncertainty metric.</p>
</blockquote>
<ul>
<li><a href="https://arxiv.org/abs/1506.02142">Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning</a></li>
</ul>
<blockquote class="blockquote">
<p>…we develop a new theoretical framework casting dropout training in deep neural networks (NNs) as approximate Bayesian inference in deep Gaussian processes. A direct result of this theory gives us tools to model uncertainty with dropout NNs</p>
</blockquote>
<ol start="3" type="1">
<li>This article wasn’t really shorter than if I had done multiple papers, less deeply. I’ll have to practice at that, not least because it’s time-consuming, but information is valuable. How does Adrian Colyer do this every <em>day</em>?</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/way-off-policy-batch-drl/</guid>
  <pubDate>Sun, 14 Jul 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/way-off-policy-batch-drl.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>A New Series arXiv Sampler</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/a-new-series-arxiv-sampler/</link>
  <description><![CDATA[ 





<section id="new-series" class="level1">
<h1>New series</h1>
<p>This post begins a weekly series highlighting one or more RL papers in the previous week’s cs.AI arXiv stream that caught my eye (making no guarantees about the correlation between what catches my eye and what ultimately turns out to be useful, important, etc). I’ll be prioritizing sustainability over most other factors, but I do hope to show you some code from time to time.</p>
<p>I read these papers to differing degrees as I have time, so there will likely be some variability in descriptive volume. However, I do pledge to make only justified statements about them so far as I know, and I welcome errata in the comments. I’m still experimenting with the format and voice, so please leave me feedback early and often to influence the series.</p>
</section>
<section id="this-week" class="level1">
<h1>This week</h1>
<p>All of this week’s papers piqued my interest because of the sample-efficiency problem in modern DRL. Reinforcement learning algorithms need to interact with the environment quite a bit before they become good at a task, and anything that can shorten this time is of interest. My group is currently working on a learning task with a very low sample rate, so we are actively on the hunt for anything that improves sample efficiency.</p>
<ul>
<li><a href="https://arxiv.org/abs/1906.12266">Growing Action Spaces</a>, by Farquhar et al.&nbsp;at Oxford and Facebook AI Research.</li>
<li><a href="https://arxiv.org/abs/1906.10187v2">Learning to Interactively Learn and Assist</a>, by Woodward et al.&nbsp;at Google Brain.</li>
<li><a href="https://arxiv.org/abs/1902.06007v2">ProLoNets: Neural-encoding Human Experts’ Domain Knowledge to Warm Start Reinforcement Learning</a>, by Silva et al.&nbsp;at Georgia Institute of Technology.</li>
</ul>
<section id="growing-action-spaces" class="level2">
<h2 class="anchored" data-anchor-id="growing-action-spaces">Growing Action Spaces</h2>
<p>Growing Action Spaces proposes a form of “curriculum learning”, where a more complex task is broken down into a sequence of simpler tasks, sometimes by humans, sometimes automatically. In this case, the authors improved the learning speed of their agent by initially giving it fewer actions to work with, training for a while, and then alternating between giving it more actions to work with and training.</p>
<p>Interestingly, they were working in Starcraft, which is a real-time strategy (RTS) game, where you have to control multiple units simultaneously in a coordinated fashion to achieve some goal. Thus, in their domain, the size of the action space didn’t just come from continuity or a really large discrete action space, but from the fact that the actions they were capable of taking were <em>combinatorial</em>. That is, they had to train an agent to take actions from a space including any combination of primitive actions, as well as any combinations of units; a daunting task.</p>
<p>Their solution is brilliant, and highly general: The authors broke the action space up into a hierarchy of action spaces by grouping units, and requiring that the same action be taken by all units within the same group. Then as training progressed, more groups were allowed to act independently. This resulted in a tractable problem at each stage of training, and overall high-performance policies that would have been prohibitively complex with conventional DRL algorithms.</p>
<p>If you or I want to apply this method to our own problems, the key requirement is to come up with a suitable way of breaking large action spaces into hierarchies of progressively smaller ones.</p>
</section>
<section id="learning-to-interactively-learn-and-assist" class="level2">
<h2 class="anchored" data-anchor-id="learning-to-interactively-learn-and-assist">Learning to Interactively Learn and Assist</h2>
<p>Reinforcement learning typically depends on a sparse reward signal and random exploration, both of which contribute to poor sample efficiency in modern algorithms. One method of improving sample efficiency and solving the exploration problem is imitation learning, where the agent is pre-trained to mimic expert behavior. However, expert demonstrations are expensive, and it’s often difficult to know how much and of what kind will suffice. These are the problems Learning to Interactively Learn and Assist attempts to solve by proposing a different paradigm entirely: without explicit demonstrations or reward function.</p>
<p>The goal is for an agent and a “principal” (say, a human) to learn to work together to accomplish the principal’s purpose. The agent takes its cues from the principal’s behavior, and acts helpfully. This requires prior understanding, both of the environment and of what constitutes communication from the principal.</p>
<p>To get to this point, the authors trained an agent jointly with a “human surrogate” principal on a variety of tasks in the same environment. Each time, the principal knows the task (as part of its observation input), and the agent does not. They receive a joint reward at the end of the episode.</p>
<blockquote class="blockquote">
<p>By informing the principal of the current task and withholding rewards and gradient updates until the end of each task, the agents are encouraged to emerge interactive learning behaviors in order to inform the assistant of the task and allow them to contribute to the joint reward.</p>
</blockquote>
<p>Prior domain knowledge required to jointly accomplish a given task is trained into the agent ahead of time this way, along with the methods of communication. Actions and observations are restricted to the environment, so that later the principal may be replaced with a human.</p>
</section>
<section id="prolonets-neural-encoding-human-experts-domain-knowledge-to-warm-start-reinforcement-learning" class="level2">
<h2 class="anchored" data-anchor-id="prolonets-neural-encoding-human-experts-domain-knowledge-to-warm-start-reinforcement-learning">ProLoNets: Neural-encoding Human Experts’ Domain Knowledge to Warm Start Reinforcement Learning</h2>
<p>ProLoNets stands for “Propositional Logic Nets”, which are a neural network architecture and method of initialization that allows a domain expert to encode initial behavior for a DRL agent in the form of propositional logic.</p>
<p>To give you the flavor:</p>
<blockquote class="blockquote">
<p>To illustrate this more practically, we consider the simplest case of a cart pole ProLoNet with a single decision node. Assume we have solicited the following from a domain expert: “If the cart’s <img src="https://latex.codecogs.com/png.latex?x"> position is right of center, move left; otherwise, move right,” and that they indicate <code>x_position</code> is the first input feature and that the center is at 0. We therefore initialize our primary node <img src="https://latex.codecogs.com/png.latex?D_0"> with <img src="https://latex.codecogs.com/png.latex?w_0=%5B1,0,0,0%5D"> and <img src="https://latex.codecogs.com/png.latex?c_0=0">. We then specify <img src="https://latex.codecogs.com/png.latex?l_0"> to be a new leaf with a prior of <img src="https://latex.codecogs.com/png.latex?%5B1,0%5D">. Finally, we set the path to <img src="https://latex.codecogs.com/png.latex?l_0"> to be <img src="https://latex.codecogs.com/png.latex?D_0"> and the path <img src="https://latex.codecogs.com/png.latex?l_1"> to be <img src="https://latex.codecogs.com/png.latex?(1-D_0)">. Consequently for each state, the probability distribution over the agent’s two actions is a softmax over <img src="https://latex.codecogs.com/png.latex?(D_0*l_0+(1-D_0)*l_1)"></p>
</blockquote>
<p>I’ve barely skimmed this paper so I don’t know what each of the components means, but I gather that a human-authored decision tree can be translated directly into a correctly-initialized neural network architecture, and an actor-critic algorithm takes over from there to improve beyond the human expert’s baseline.</p>
<p>Something else that caught my eye:</p>
<blockquote class="blockquote">
<p>While our initialized ProLoNets are able to follow expert strategies immediately, they may lack expressive capacity to learn more optimal policies once they are deployed into a domain. … To enable the ProLoNet architecture to continue to grow beyond its initial definition, we introduce a dynamic deepening procedure.</p>
</blockquote>
<blockquote class="blockquote">
<p>Upon initialization, a ProLoNet agent maintains two copies of its actor: the shallower, unaltered initialized version and a deeper version, in which each leaf is transformed into a randomly initialized node with two new randomly initialized leaves. As the agent interacts with its environment, it relies on the shallower networks to generate actions and value predictions and to gather experience, After each episode, our off-policy update is run over the shallower and deeper networks. Finally, after the off-policy updates, the agent compares the entropy of the shallower actor’s leaves to the entropy of the deeper actor’s leaves and selectively deepens when the leaves of the deeper actor are less uniform than those of the shallower actor. We find that this dynamic deepening improves stability and ameliorates policy degradation.</p>
</blockquote>
<p>This strikes me as the beginning of the future, where neural network architecture is learned and adjusted dynamically alongside the network parameters.</p>
</section>
</section>
<section id="parting-thoughts" class="level1">
<h1>Parting thoughts</h1>
<ol type="1">
<li>I’m extremely pleased to have finally gotten this off the ground. Please comment on anything and everything, and we’ll drive this thing together.</li>
<li>Growing Action Spaces is immediately relevant to my group, since in the medium-term, we intend to increase our action spaces combinatorially, and will inherit all of the trouble this brings. More on this another time.</li>
<li>I wonder how often in complex real environments the “Learning to Interactively Learn and Assist” agents will learn to communicate in a way that humans find unintuitive. Since the quickest way to communicate involves some compression, would we need to add some term representing human understandability? How best to do this?</li>
<li>“Learning to Interactively Learn and Assist” seems like a relevant paper for AI safety, though as far as I could tell in my quick read, it wasn’t billed that way. If we train agents that don’t have goals of their own necessarily, but take their cues from us in real time, are we safer than if we attempted to craft the perfect reward function, or demonstrated our desires in a one-and-done fashion?</li>
<li>I’ve gotta actually read the ProLoNets paper. There was even more to it than I highlighted, and they included an ablation study which will likely tell me if I can incorporate their concepts piecemeal into my own work.</li>
</ol>


</section>

 ]]></description>
  <category>arXiv highlights</category>
  <guid>https://computable.ai/posts/a-new-series-arxiv-sampler/</guid>
  <pubDate>Sun, 07 Jul 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/a-new-series-arxiv-sampler.svg" medium="image" type="image/svg+xml"/>
</item>
<item>
  <title>Boltzmann Machines: Differentiation Work</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/boltzmann-machines-differentiation-work/</link>
  <description><![CDATA[ 





<p>I recently read <a href="https://theneural.wordpress.com/2011/07/08/the-miracle-of-the-boltzmann-machine/">The Miracle of the Boltzmann Machine</a>, and it’s so compelling that I’ve been thinking about it ever since. I intend to write much more on Boltzmann Machines in the future, but here I’m just going to show my work differentiating the objective function.</p>
<section id="given" class="level3">
<h3 class="anchored" data-anchor-id="given">Given</h3>
<ol type="1">
<li>Objective function <img src="https://latex.codecogs.com/png.latex?L(W)%20:=%20%5Cmathbb%7BE%7D_%7BD(V)%7D%20%5Blog%20P(V)%5D"></li>
<li>and probability of a given BM state <img src="https://latex.codecogs.com/png.latex?X=(V,H)"> <img src="https://latex.codecogs.com/png.latex?P(X)%20:=%20P(V,H)%20:=%20%7Be%5E%7BX%5ETWX/2%7D%5Cover%20%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D%7D"> <img src="https://latex.codecogs.com/png.latex?P(V)%20:=%20%5Csum_H%20P(V,H)%20=%20%5Cfrac%7B%5Csum_H%20e%5E%7BX%5ETWX/2%7D%7D%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D"> where <img src="https://latex.codecogs.com/png.latex?W"> is the BM transition matrix, assuming <img src="https://latex.codecogs.com/png.latex?w_%7Bij%7D=w_%7Bji%7D"></li>
</ol>
</section>
<section id="want-to-show" class="level3">
<h3 class="anchored" data-anchor-id="want-to-show">Want to show</h3>
<p><img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20L%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%20%5Cmathbb%7BE%7D_%7BD(V)P(H%7CV)%7D%5Bx_ix_j%5D-%5Cmathbb%7BE%7D_%7BP(V,H)%7D%5Bx_ix_j%5D"></p>
</section>
<section id="proof" class="level3">
<h3 class="anchored" data-anchor-id="proof">Proof</h3>
<ol type="1">
<li>Definition of expected value <img src="https://latex.codecogs.com/png.latex?L(W)=%5Cmathbb%7BE%7D_%7BD(V)%7D%20%5B%5Clog%20P(V)%5D%20=%20%5Csum_V%20D(V)%5Clog%20P(V)"></li>
<li>Let <img src="https://latex.codecogs.com/png.latex?f%20=%20logP(V)"> <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20L%7D%7B%5Cpartial%20f%7D%20=%20%5Csum_V%20D(V)%5Cfrac%7B%5Cpartial%20f%7D%7B%5Cpartial%20w_%7Bij%7D%7D"></li>
<li>Chain rule <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20f%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%20%7B%5Cfrac%7B%5Cpartial%20P(V)%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20%5Cover%20P(V)%7D"></li>
<li>Expand <img src="https://latex.codecogs.com/png.latex?P(V)"> <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20P(V)%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%20%5Cfrac%7B%5Cpartial%7D%7B%5Cpartial%20w_%7Bij%7D%7D%5Cleft%5B%5Csum_H%20P(V,H)%5Cright%5D%20=%20%5Cfrac%7B%5Cpartial%7D%7B%5Cpartial%20w_%7Bij%7D%7D%5Cleft%5B%5Csum_H%20%7Be%5E%7BX%5ETWX/2%7D%5Cover%20%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D%7D%5Cright%5D%20=%20%5Csum_H%20%5Cfrac%7B%5Cpartial%7D%7B%5Cpartial%20w_%7Bij%7D%7D%5Cleft%5B%7Be%5E%7BX%5ETWX/2%7D%5Cover%20%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D%7D%5Cright%5D"></li>
<li>Quotient rule <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20P(V)%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%5Csum_H%20%5Cfrac%7B%5Cfrac%7B%5Cpartial%7D%7B%5Cpartial%20w_%7Bij%7D%7D%5Cleft%5Be%5E%7BX%5ETWX/2%7D%5Cright%5D%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D-e%5E%7BX%5ETWX/2%7D%20%5Cfrac%7B%5Cpartial%7D%7B%5Cpartial%20w_%7Bij%7D%7D%5Cleft%5B%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D%5Cright%5D%7D%7B%5Cleft(%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D%5Cright)%5E2%7D"></li>
<li>Chain rule, and notice <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%7D%7B%5Cpartial%20w_%7Bij%7D%7D%5Cleft%5BW%5Cright%5D"> is <img src="https://latex.codecogs.com/png.latex?0"> everywhere except <img src="https://latex.codecogs.com/png.latex?w_%7Bij%7D">, so <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%7D%7B%5Cpartial%20w_%7Bij%7D%7D%5Cleft%5Be%5E%7BX%5ETWX/2%7D%5Cright%5D%20=%20%5Cfrac%7B%5Cpartial%7D%7B%5Cpartial%20w_%7Bij%7D%7D%5Cleft%5BX%5ETWX/2%5Cright%5D%20e%5E%7BX%5ETWX/2%7D%20=%20x_ix_je%5E%7BX%5ETWX/2%7D"></li>
<li>So #5 becomes <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20P(V)%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%20%5Csum_H%20%5Cfrac%7Bx_ix_je%5E%7BX%5ETWX/2%7D%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D-e%5E%7BX%5ETWX/2%7D%20%5Csum_%7BX'%7Dx'_ix'_je%5E%7BX'%5ETWX'/2%7D%7D%7B%5Cleft(%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D%5Cright)%5E2%7D"></li>
<li>Separating terms <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20P(V)%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%20%5Csum_H%5Cleft%5B%5Cfrac%7Bx_ix_je%5E%7BX%5ETWX/2%7D%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D%7D%7B%5Cleft(%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D%5Cright)%5E2%7D%5Cright%5D-%5Csum_H%5Cleft%5B%5Cfrac%7Be%5E%7BX%5ETWX/2%7D%20%5Csum_%7BX'%7Dx'_ix'_je%5E%7BX'%5ETWX'/2%7D%7D%7B%5Cleft(%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D%5Cright)%5E2%7D%5Cright%5D"></li>
<li>Cancelling and moving factors outside sums <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20P(V)%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%20%5Csum_H%5Cleft%5B%5Cfrac%7Bx_ix_je%5E%7BX%5ETWX/2%7D%7D%7B%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D%7D%5Cright%5D-%5Cfrac%7B%5Csum_H%5Cleft%5Be%5E%7BX%5ETWX/2%7D%5Cright%5D%20%5Csum_%7BX'%7Dx'_ix'_je%5E%7BX'%5ETWX'/2%7D%7D%7B%5Cleft(%7B%5Csum_%7BX'%7D%20e%5E%7BX'%5ETWX'/2%7D%7D%5Cright)%5E2%7D"></li>
<li>Definition of <img src="https://latex.codecogs.com/png.latex?P(V,H)"> and <img src="https://latex.codecogs.com/png.latex?P(V)"> <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20P(V)%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%20%5Csum_H%5Cleft%5Bx_ix_jP(V,H)%5Cright%5D-P(V)%20%5Csum_%7BX'%7D%5Cleft%5Bx'_ix'_jP(V',H')%5Cright%5D"></li>
<li>Substituting #10 into #3 and #3 into #2 we have <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20L%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%20%5Csum_VD(V)%5Cleft%5B%5Cfrac%7B%5Csum_H%5Cleft%5Bx_ix_jP(V,H)%5Cright%5D-P(V)%20%5Csum_%7BX'%7D%5Cleft%5Bx'_ix'_jP(V',H')%5Cright%5D%7D%7BP(V)%7D%5Cright%5D"></li>
<li>Separating into two terms <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20L%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%20%5Csum_V%5Cleft%5BD(V)%5Csum_H%5Cleft%5B%5Cfrac%7Bx_ix_jP(V,H)%7D%7BP(V)%7D%5Cright%5D%5Cright%5D-%5Csum_V%5Cleft%5BD(V)P(V)%5Csum_%7BX'%7D%5Cleft%5Bx'_ix'_jP(V',H')%5Cright%5D%5Cright%5D"></li>
<li>Definition of conditional probability <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20L%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%20%5Csum_V%5Csum_H%5Cleft%5Bx_ix_jD(V)P(H%7CV)%5Cright%5D-%5Csum_VD(V)%5Csum_%7BX'%7D%5Cleft%5Bx'_ix'_jP(V',H')%5Cright%5D"></li>
<li><img src="https://latex.codecogs.com/png.latex?%5Csum_VD(V)=1">, combining sums, and <img src="https://latex.codecogs.com/png.latex?X=(V,H)"> <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20L%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%5Csum_%7B(V,H)%7D%5Cleft%5Bx_ix_jD(V)P(H%7CV)%5Cright%5D-%5Csum_%7B(V',H')%7D%5Cleft%5Bx'_ix'_jP(V',H')%5Cright%5D"></li>
<li>Definition of expected value <img src="https://latex.codecogs.com/png.latex?%5Cfrac%7B%5Cpartial%20L%7D%7B%5Cpartial%20w_%7Bij%7D%7D%20=%20%5Cmathbb%7BE%7D_%7BD(V)P(H%7CV)%7D%5Bx_ix_j%5D-%5Cmathbb%7BE%7D_%7BP(V,H)%7D%5Bx_ix_j%5D"> <img src="https://latex.codecogs.com/png.latex?%5Csquare"></li>
</ol>


</section>

 ]]></description>
  <category>Math</category>
  <guid>https://computable.ai/posts/boltzmann-machines-differentiation-work/</guid>
  <pubDate>Sun, 10 Mar 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/boltzmannexample.png" medium="image" type="image/png" height="137" width="144"/>
</item>
<item>
  <title>Inaugural Post</title>
  <dc:creator>Daniel Cox</dc:creator>
  <link>https://computable.ai/posts/inaugural-post/</link>
  <description><![CDATA[ 





<p>This post begins the Computable AI blog, a machine intelligence blog from a handful of DRL practitioners, intended to crystalize, internalize, share, and explain.</p>
<p>I found few beginner resources for DRL when I began, and since I have a passion for teaching, this seemed a likely area in which to make a dent.</p>
<p>I also serve as the “Director of Applied Sciences” for a startup software company, and the AI team must occasionally indoctrinate new members. This provides us with a convenient target audience, as well as an expanding pool of co-authors.</p>
<p>Finally, my own education in DRL is incomplete, so this will serve partly as a record of my own journey.</p>
<p>I hope it helps you.</p>



 ]]></description>
  <category>Miscellany</category>
  <guid>https://computable.ai/posts/inaugural-post/</guid>
  <pubDate>Sat, 16 Feb 2019 00:00:00 GMT</pubDate>
  <media:content url="https://computable.ai/static/images/thumbs/inaugural-post.svg" medium="image" type="image/svg+xml"/>
</item>
</channel>
</rss>
