I had a really long post written out, and then decided to check your link. It turns out that I what I was suggesting was the second of the three -- the pure Client/Server setup. But I'm suggesting, from what you're asking, that you want the Client/Server setup with Client-side prediction (off the list, the best of the three, so it'd be stupid for me to argue against it

).
The first thought off the top of my head would be something rudimentary like so:
Code: Select all
void MoveObject(object, previousTransform, velocity, angularVelocity, deltaT)
{
object.setWorldTransformation(previousTransform);
object.setLinearVelocity(velocity);
object.setAngularVelocity(velocity);
world.stepSimulation(deltaT);
}
None of the code above is meant to be directly translated -- just barely above pseudocode. But if you set up where it was at the last transformation, then set the velocity and the angular velocity, you could step the world by however long it's been since, and then get the new information for where it is at. Since I'm assuming that you'd want to have the object move more often than once every time the server spoke, you could have two Bullet worlds
on the client -- one for the client itself, and one pseudo-world for the server (on the client). The client world would be what you'd keep everything in, and would be what you'd be using primarily. But for every object in the client world, you'd have one in the server world (unless I say otherwise, "server world" will mean this secondary client-side world). And when you get information from the server machine, you could update the objects appropriately using the above function in the server world, and then step the server world by the deltaT. With that all done, you'd have updated positions and information that you could just copy from the server world into the client world (easily enough done through a MotionState). If the server world and the client world information both match up, then you're golden; if they don't, then you'd need to correct appropriately (you'd want to trust the server world over the client world, because you always want the client to reflect what's going on in the server machine).
As for collision, I might suggest leaving this up to the server machine itself. Unless you're writing a P2P program, it means you'd have a server machine that you get to choose, and generally server machines run much better than the average client computer. Plus, with a server machine, you could ignore any rendering and allow it to work with "pure data" -- what I'm getting at is that it'd be much faster, and since every client is already speaking to it, you wouldn't need to have it pass information from client A to client B about what client A is doing in client A's world; the server machine would just tell client B how client B should respond to what client A did. Because the server would run faster than the clients, I would have collision done on the server, and just report any collisions back to the client. Again, I could be wrong, so my word isn't gospel, but that seems like it would be faster and less prone to client-side mistakes or mismatches (or, if you're worried about it, cheating

).