<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en-gb">
	<link rel="self" type="application/atom+xml" href="https://pybullet.org/Bullet/phpBB3/app.php/feed/topic/1249" />

	<title>Real-Time Physics Simulation Forum</title>
	
	<link href="https://pybullet.org/Bullet/phpBB3/index.php" />
	<updated>2007-06-26T08:32:19+00:00</updated>

	<author><name><![CDATA[Real-Time Physics Simulation Forum]]></name></author>
	<id>https://pybullet.org/Bullet/phpBB3/app.php/feed/topic/1249</id>

		<entry>
		<author><name><![CDATA[Dirk Gregorius]]></name></author>
		<updated>2007-06-26T08:32:19+00:00</updated>

		<published>2007-06-26T08:32:19+00:00</published>
		<id>https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4596#p4596</id>
		<link href="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4596#p4596"/>
		<title type="html"><![CDATA[SIMD Vector3 Math]]></title>

		
		<content type="html" xml:base="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4596#p4596"><![CDATA[
May I ask on which experience you base all your statements here? Did you test this in a MLOC project or did you write some simple testbed like this:<br><div class="codebox"><p>Code: </p><pre><code>int main( void ){BEGIN_PROFILE( "NON_SSE_CROSS" );v = cross( v1, v2 );END_PROFILE();BEGIN_PROFILE( "SSE_CROSS" ); v = cross_sse( v1, v2 );END_PROFILE();if ( time_non_sse &lt; time_sse )  printf( "I did a pretty good job!" );return 0;}</code></pre></div><p>Statistics: Posted by <a href="https://pybullet.org/Bullet/phpBB3/memberlist.php?mode=viewprofile&amp;u=14">Dirk Gregorius</a> — Tue Jun 26, 2007 8:32 am</p><hr />
]]></content>
	</entry>
		<entry>
		<author><name><![CDATA[Ernegien]]></name></author>
		<updated>2007-06-25T18:23:23+00:00</updated>

		<published>2007-06-25T18:23:23+00:00</published>
		<id>https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4588#p4588</id>
		<link href="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4588#p4588"/>
		<title type="html"><![CDATA[SIMD Vector3 Math]]></title>

		
		<content type="html" xml:base="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4588#p4588"><![CDATA[
I never made such claims, other than providing a few examples of SIMD implementation in vector math operations, which do offer a small performance increase (the Normalize() and Cross() methods in particular) over other standard methods.<br><br>The article brings up some good points, but intrinsics aren't any better than what I'm currently doing, aside from eliminating a single procedure call, which will most likely get canceled out by poor compiler optimisation anyways.  That and maybe a few other dependencies that can be easily avoided by switching up the opcode orders or register assignments in some of my functions, but for the most part, I believe I did a fairly good job doing it by hand...feel free to correct me if I'm wrong though ;P<br><br>And again, I'm not here to argue that SIMD solves everything...but these general concepts can be easily used to help speed up certain segments of code that have lots of calculations and take more time to execute.<p>Statistics: Posted by <a href="https://pybullet.org/Bullet/phpBB3/memberlist.php?mode=viewprofile&amp;u=1802">Ernegien</a> — Mon Jun 25, 2007 6:23 pm</p><hr />
]]></content>
	</entry>
		<entry>
		<author><name><![CDATA[Dirk Gregorius]]></name></author>
		<updated>2007-06-25T08:46:26+00:00</updated>

		<published>2007-06-25T08:46:26+00:00</published>
		<id>https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4579#p4579</id>
		<link href="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4579#p4579"/>
		<title type="html"><![CDATA[SIMD Vector3 Math]]></title>

		
		<content type="html" xml:base="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4579#p4579"><![CDATA[
<blockquote class="uncited"><div>From the text:<br><br>On programming forums, every once in a while someone appears who wants to optimize his program by replacing his current 3D vector class by one that uses SSE opcodes, in the hope he'll make his program run 4 times as fast.</div></blockquote>What is new is that now people who state that they are fairly new to SIMD operations start giving advices on programming forums. I wonder when we get the first post here that states using C++ is idiotic and that we should use Java or C# instead "since it is only 5% slower"....<p>Statistics: Posted by <a href="https://pybullet.org/Bullet/phpBB3/memberlist.php?mode=viewprofile&amp;u=14">Dirk Gregorius</a> — Mon Jun 25, 2007 8:46 am</p><hr />
]]></content>
	</entry>
		<entry>
		<author><name><![CDATA[Pierre]]></name></author>
		<updated>2007-06-25T07:32:51+00:00</updated>

		<published>2007-06-25T07:32:51+00:00</published>
		<id>https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4578#p4578</id>
		<link href="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4578#p4578"/>
		<title type="html"><![CDATA[SIMD Vector3 Math]]></title>

		
		<content type="html" xml:base="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4578#p4578"><![CDATA[
<a href="http://www.farbrausch.de/~fg/articles/ubiquitous_sse_vector.html" class="postlink">http://www.farbrausch.de/~fg/articles/u ... ector.html</a><br><br>*yawn*<p>Statistics: Posted by <a href="https://pybullet.org/Bullet/phpBB3/memberlist.php?mode=viewprofile&amp;u=76">Pierre</a> — Mon Jun 25, 2007 7:32 am</p><hr />
]]></content>
	</entry>
		<entry>
		<author><name><![CDATA[Ernegien]]></name></author>
		<updated>2007-06-24T18:54:16+00:00</updated>

		<published>2007-06-24T18:54:16+00:00</published>
		<id>https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4573#p4573</id>
		<link href="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4573#p4573"/>
		<title type="html"><![CDATA[SIMD Vector3 Math]]></title>

		
		<content type="html" xml:base="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4573#p4573"><![CDATA[
Interesting article <img class="smilies" src="https://pybullet.org/Bullet/phpBB3/images/smilies/icon_smile.gif" width="15" height="15" alt=":)" title="Smile">  Anyways, I guess the purpose of my post is to inform people about the benefits SIMD has to offer.  Sure, you won't notice any solid performance gains if you're only executing these every once in a while...but when you are calling them on a continuous basis, there will be a noticeable difference (like any other optimised piece of code).  Even if its only a few thousand more executions per second, it's still worth it in my opinion.  Now I'm not saying to go crazy and convert everything to SIMD, because for the most part you're right, there's no need.  But little things like avoiding a square root when normalizing a vector are too good to pass up <img class="smilies" src="https://pybullet.org/Bullet/phpBB3/images/smilies/icon_razz.gif" width="15" height="15" alt=":P" title="Razz"><br><br>I'll update my post above with some more I've managed to convert, all of which offer considerable performance gains over their C++ counterpart.<p>Statistics: Posted by <a href="https://pybullet.org/Bullet/phpBB3/memberlist.php?mode=viewprofile&amp;u=1802">Ernegien</a> — Sun Jun 24, 2007 6:54 pm</p><hr />
]]></content>
	</entry>
		<entry>
		<author><name><![CDATA[Dirk Gregorius]]></name></author>
		<updated>2007-06-24T17:22:54+00:00</updated>

		<published>2007-06-24T17:22:54+00:00</published>
		<id>https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4572#p4572</id>
		<link href="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4572#p4572"/>
		<title type="html"><![CDATA[SIMD Vector3 Math]]></title>

		
		<content type="html" xml:base="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4572#p4572"><![CDATA[
Did you experience these improvements in a real application or just by measuring a SIMD cross product vs. a usual implementation? My experience is that a SIMD implementation as you suggest here brings pretty much nothing on the PC. Even worth it can actually slow down things. On the other hand using SIMD for time critical code can bring huge improvements, e.g. <br><br><a href="http://www.intel.com/cd/ids/developer/asmo-na/eng/293451.htm?prn=Y" class="postlink">http://www.intel.com/cd/ids/developer/a ... .htm?prn=Y</a><br><br>Implementing a SIMD math library looks like a trivial straight forward thing, but actually it isn't. Also there is a huge difference between PC and PPC.<br><br>So again, what is the sense of this post? Do you want to teach people writing assembly code?<p>Statistics: Posted by <a href="https://pybullet.org/Bullet/phpBB3/memberlist.php?mode=viewprofile&amp;u=14">Dirk Gregorius</a> — Sun Jun 24, 2007 5:22 pm</p><hr />
]]></content>
	</entry>
		<entry>
		<author><name><![CDATA[Ernegien]]></name></author>
		<updated>2007-06-24T16:05:01+00:00</updated>

		<published>2007-06-24T16:05:01+00:00</published>
		<id>https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4571#p4571</id>
		<link href="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4571#p4571"/>
		<title type="html"><![CDATA[SIMD Vector3 Math]]></title>

		
		<content type="html" xml:base="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4571#p4571"><![CDATA[
Huge, no.  The compiler does a fairly decent job...but I've experienced a 50-100% increase in speed in methods that could directly benefit from such optimisations.  It isn't much, but every little tick counts ;P<p>Statistics: Posted by <a href="https://pybullet.org/Bullet/phpBB3/memberlist.php?mode=viewprofile&amp;u=1802">Ernegien</a> — Sun Jun 24, 2007 4:05 pm</p><hr />
]]></content>
	</entry>
		<entry>
		<author><name><![CDATA[Dirk Gregorius]]></name></author>
		<updated>2007-06-24T10:27:10+00:00</updated>

		<published>2007-06-24T10:27:10+00:00</published>
		<id>https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4568#p4568</id>
		<link href="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4568#p4568"/>
		<title type="html"><![CDATA[SIMD Vector3 Math]]></title>

		
		<content type="html" xml:base="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4568#p4568"><![CDATA[
What do you want to show with this? Do you expect that replacing your math library with as a generic SIMD implementation will give huge performance improvements?<p>Statistics: Posted by <a href="https://pybullet.org/Bullet/phpBB3/memberlist.php?mode=viewprofile&amp;u=14">Dirk Gregorius</a> — Sun Jun 24, 2007 10:27 am</p><hr />
]]></content>
	</entry>
		<entry>
		<author><name><![CDATA[Ernegien]]></name></author>
		<updated>2007-06-24T18:56:26+00:00</updated>

		<published>2007-06-23T17:41:31+00:00</published>
		<id>https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4567#p4567</id>
		<link href="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4567#p4567"/>
		<title type="html"><![CDATA[SIMD Vector3 Math]]></title>

		
		<content type="html" xml:base="https://pybullet.org/Bullet/phpBB3/viewtopic.php?p=4567#p4567"><![CDATA[
For those that wish to see for themselves the performance gains associated with SIMD operations, here's a few functions to benchmark against the other standard ones.  I'm still fairly new to SIMD operations in assembly, but if anyone needs help with some of the math optimisation, feel free to ask <img class="smilies" src="https://pybullet.org/Bullet/phpBB3/images/smilies/icon_wink.gif" width="15" height="15" alt=";)" title="Wink"><br><div class="codebox"><p>Code: </p><pre><code>#pragma oncestatic const float SZero = 0; __declspec(align(16)) struct Vector3D{#pragma region Constructorpublic: Vector3D(){X = 0;Y = 0;Z = 0;W = 0;}public: Vector3D(float x, float y, float z){X = x;Y = y;Z = z;W = 0;}#pragma endregion#pragma region Destructorpublic: ~Vector3D(void){}#pragma endregion#pragma region Fieldspublic: float X;public: float Y;public: float Z;private: float W;#pragma endregion#pragma region Properties#pragma endregion#pragma region Operator Overloads#pragma endregion#pragma region Methodspublic: void NormalizePrecise(){_asm{//get lengthmoveax, thismovdxmm0, dword ptr ds:[eax]movdxmm1, dword ptr ds:[eax + 4]movdxmm2, dword ptr ds:[eax + 8]mulssxmm0, xmm0mulssxmm1, xmm1mulssxmm2, xmm2addssxmm0, xmm1addssxmm0, xmm2sqrtssxmm0, xmm0ucomissxmm0, SZerojeDone//duplicate length across registerpshufdxmm0, xmm0, 0//divide by lengthmovdqaxmm1, xmmword ptr ds:[eax]divpsxmm1, xmm0//store resultmovdqaxmmword ptr ds:[eax], xmm1Done:}}public: void Normalize(){_asm{//get recipricol lengthmoveax, thismovdxmm0, dword ptr ds:[eax]movdxmm1, dword ptr ds:[eax + 4]movdxmm2, dword ptr ds:[eax + 8]mulssxmm0, xmm0mulssxmm1, xmm1mulssxmm2, xmm2addssxmm0, xmm1addssxmm0, xmm2ucomissxmm0, SZerojeDonersqrtssxmm0, xmm0//duplicate length across registerpshufdxmm0, xmm0, 0//multiply by recipricol length (divide by length)movdqaxmm1, xmmword ptr ds:[eax]mulpsxmm1, xmm0//store resultmovdqaxmmword ptr ds:[eax], xmm1Done:}}public: void Absolute(){_asm{moveax, thisanddword ptr ds:[eax], 07FFFFFFFhanddword ptr ds:[eax + 4], 07FFFFFFFhanddword ptr ds:[eax + 8], 07FFFFFFFh}}public: void Maximize(const Vector3D &amp;v1){_asm{//get parameter informationmoveax, v1movdqaxmm0, xmmword ptr ds:[eax]moveax, this//compute maximummaxpsxmm0, xmmword ptr ds:[eax]//store resultmovdqaxmmword ptr ds:[eax], xmm0}}public: void Minimize(const Vector3D &amp;v1){_asm{//get parameter informationmoveax, v1movdqaxmm0, xmmword ptr ds:[eax]moveax, this//compute minimumminpsxmm0, xmmword ptr ds:[eax]//store resultmovdqaxmmword ptr ds:[eax], xmm0}}public: void Cross(const Vector3D &amp;v1){_asm{//get parameter informationmoveax, v1movdqaxmm0, xmmword ptr ds:[eax]moveax, thismovdqaxmm1, xmmword ptr ds:[eax]//align vectors to be multipliedpshufdxmm2, xmm0, 11001001b//(v1.Y, v1.Z, v1.X)pshufdxmm4, xmm1,11001001b//(Y, Z, X)pshufdxmm3, xmm0, 11010010b//(v1.Z, v1.X, v1.Y)pshufdxmm5, xmm1,11010010b//(Z, X, Y)//perform cross-productmulpsxmm4, xmm3mulpsxmm5, xmm2subpsxmm4, xmm5//store resultmovdqaxmmword ptr ds:[eax], xmm4}}public: void Lerp(const Vector3D &amp;v1, float interpolator){_asm{//get parameter informationmoveax, v1movdqaxmm0, xmmword ptr ds:[eax]moveax, thismovdqaxmm1, xmmword ptr ds:[eax]//duplicate interpolator across registerpshufdxmm2, interpolator, 0//interpolatesubpsxmm0, xmm1mulpsxmm0, xmm2addpsxmm0, xmm1//store resultmovdqaxmmword ptr ds:[eax], xmm0}}#pragma endregion};</code></pre></div>Please note that this requires your vector to be alligned on a 16-byte boundary.  Also, this was compiled as a win32 project using Visual Studios 2005.  You may need to modify the assembly a bit depending on which compiler you use.<p>Statistics: Posted by <a href="https://pybullet.org/Bullet/phpBB3/memberlist.php?mode=viewprofile&amp;u=1802">Ernegien</a> — Sat Jun 23, 2007 5:41 pm</p><hr />
]]></content>
	</entry>
	</feed>
