<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en-gb">
	<link rel="self" type="application/atom+xml" href="https://pybullet.org/Bullet/phpBB3/app.php/feed/topic/13005" />

	<title>Real-Time Physics Simulation Forum</title>
	
	<link href="https://pybullet.org/Bullet/phpBB3/index.php" />
	<updated>2020-06-22T21:28:20+00:00</updated>

	<author><name><![CDATA[Real-Time Physics Simulation Forum]]></name></author>
	<id>https://pybullet.org/Bullet/phpBB3/app.php/feed/topic/13005</id>

		<entry>
		<author><name><![CDATA[djb]]></name></author>
		<updated>2020-06-22T21:28:20+00:00</updated>

		<published>2020-06-22T21:28:20+00:00</published>
		<id>https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=42947#p42947</id>
		<link href="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=42947#p42947"/>
		<title type="html"><![CDATA[ARS policy understanding]]></title>

		
		<content type="html" xml:base="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=42947#p42947"><![CDATA[
Hi<br><br>I am able to train my robot in simulation using ARS, and the rewards go up, until it's doing a nice quadruped walk of about 8 metres, in 1000 time steps, per episode.<br><br>But when I stop the program, and load the latest policy file saved, the progress seems to reset.<br><br>So here are rewards before quitting:<br>                                Step: 228          Reward: 7.3702018170620045<br>                                Step: 229          Reward: 7.714461611598418<br>                                Step: 230          Reward: 7.561359695748974<br>                                Step: 231          Reward: 7.821410289581096<br>                                Step: 232          Reward: 8.063394689120104<br><br>Then I quit and start over with the policy from this run, and now the rewards are:<br><br>                                Step: 0          Reward: -0.857725762127679<br>                                Step: 1          Reward: 0.4743116543915239<br>                                Step: 2          Reward: 2.847940117990215<br>                                Step: 3          Reward: -0.8539553638529137<br>                                Step: 4          Reward: 3.0954765964267392<br>                                Step: 5          Reward: -0.5506416870964888<br>                                Step: 6          Reward: -0.626510343527105<br>                                Step: 7          Reward: 1.7169761539347284<br>                                Step: 8          Reward: 1.4009849267252874<br>                                Step: 9          Reward: 5.664102735951084<br>                                Step: 10          Reward: 0.46594051311167095<br><br><br><br>For ARS, theta is the matrix of perceptron weights between<br>  nb_inputs = env.observation_space.shape[0]<br>and<br>  nb_outputs = env.action_space.shape[0]<br><br>For my robot, with 4 legs, there's 4 actions (-1 to 1) for the 4 motors<br>And the input is 16 numbers, which is an observation:<br> (4 motor angles, 4 motor velocities, 4 motor torques, and base orientation in quaternion form)<br><br><br>Anyway, am I missing something obvious?<br><br>What might be different between carrying on from step 231 to step 232, vs starting again from step 0?<br><br>I've verified that theta being saved is the same as the theta being loaded.<br><br>Thanks<p>Statistics: Posted by <a href="https://pybullet.org/Bullet/phpBB3/memberlist.php?mode=viewprofile&amp;u=13666">djb</a> — Mon Jun 22, 2020 9:28 pm</p><hr />
]]></content>
	</entry>
	</feed>
