The pentium 4's highest bandwidth anywhere was 48GB/s through the L2 cache. The PCIe 5.0 x16 on an RTX 5060 Ti can do roughly 64 GB/s in each direction. Obviously the bandwidth in and out of the Pentium 4 are quite a bit lower than the fastest thing in the whole chip... so there's no feasible way for a pentium 4 to deliver the maximum bandwidth to a 5060 Ti.
Also, the early Pentium 4 CPUs didn't have EM64T, so the maximum addressable memory was 4GB. So some Pentium 4 CPUs are unable to use all the memory on a 5060 Ti. This almost makes it sound like one of these personified hardware components is irrational...
To copy things into GPU memory from system memory, the CPU needs an address to copy it to. There are some ways around the 4GB limit with e.g. PAE, but the limitation that a single application can only use 32 bits of memory addresses applies regardless. With CUDA there's definitely ways to fill the memory of the GPU with something - you might even be able to make limited use of it - but it's not all addressable at once from any general context, and it won't be useful for gaming purposes in any way I'm aware of.
I guess it's also worth a mention that some Pentium 4 systems didn't have PCIe - just the prior PCI, so there might need to be some kind of adapter thing in the picture for this to work out at all, and I don't know whether those do some kind of DMA, impersonation of the older standards, or something else entirely.
It probably just isn't a good idea to connect a modern GPU to a 25 year old processor.
I don't get this one. The refresh rate one makes sense. The monitor accepted a frame rate faster than it could draw. I'm not sure what this one is saying
It doesn’t have to make sense technologically it’s just a play on “how tall are you” “ how much do you weigh” (which also aren’t technical requirements for compatibility)
My best guess is that it has something to do with most render pipelines using CPU frame scheduling. Since the GPU can only process shaders, the CPU has to do e.g. input logic, reading textures from disk, etc. So in most engines, the CPU gives the GPU basic directions for each frame, and then the GPU goes off to process them while the CPU prepares the next set of directions.
If the CPU isn't fast enough, the GPU sits around waiting for instructions and it causes stuttering and FPS drops. If the GPU isn't fast enough, instructions from the CPU queue up and it causes input lag. To get around the input lag, most engines intentionally drop frames and decrease FPS (basically forcing the CPU to wait) in order to keep everything in sync.
Where VRAM comes into play is that that's where textures and other important data the GPU needs are stored. The CPU and GPU can actually both write to VRAM (though the CPU does so indirectly in most cases). Writing to VRAM is an expensive operation that benefits from parallelization, especially in the context of loading textures and cached shaders, therefore more CPU cores = better performance. Meanwhile, running out of VRAM usually causes the GPU to fall back to system RAM which is much slower, so more VRAM = better performance.
TL;DR The CPU runs the game logic and the GPU renders the result, so the GPU can only render frames as fast as the CPU can prepare them and vice-versa.
Edit: Since people seem unable to tell, I'm not a bot (obligatory fuck AI); sorry I use structured formatting tho ig ¯\_(ツ)_/¯
8 replies
The pentium 4's highest bandwidth anywhere was 48GB/s through the L2 cache. The PCIe 5.0 x16 on an RTX 5060 Ti can do roughly 64 GB/s in each direction. Obviously the bandwidth in and out of the Pentium 4 are quite a bit lower than the fastest thing in the whole chip... so there's no feasible way for a pentium 4 to deliver the maximum bandwidth to a 5060 Ti.
Also, the early Pentium 4 CPUs didn't have EM64T, so the maximum addressable memory was 4GB. So some Pentium 4 CPUs are unable to use all the memory on a 5060 Ti. This almost makes it sound like one of these personified hardware components is irrational...
Is CPU addressing GPU's memory? I thought they are just chatting over the bus.
To copy things into GPU memory from system memory, the CPU needs an address to copy it to. There are some ways around the 4GB limit with e.g. PAE, but the limitation that a single application can only use 32 bits of memory addresses applies regardless. With CUDA there's definitely ways to fill the memory of the GPU with something - you might even be able to make limited use of it - but it's not all addressable at once from any general context, and it won't be useful for gaming purposes in any way I'm aware of.
I guess it's also worth a mention that some Pentium 4 systems didn't have PCIe - just the prior PCI, so there might need to be some kind of adapter thing in the picture for this to work out at all, and I don't know whether those do some kind of DMA, impersonation of the older standards, or something else entirely.
It probably just isn't a good idea to connect a modern GPU to a 25 year old processor.
You may want to schedule an autism test with your doctors.
I don't get this one. The refresh rate one makes sense. The monitor accepted a frame rate faster than it could draw. I'm not sure what this one is saying
Don't ask a man his salary, a woman her weight, or a pentium how many cores
Or something...
It doesn’t have to make sense technologically it’s just a play on “how tall are you” “ how much do you weigh” (which also aren’t technical requirements for compatibility)
My best guess is that it has something to do with most render pipelines using CPU frame scheduling. Since the GPU can only process shaders, the CPU has to do e.g. input logic, reading textures from disk, etc. So in most engines, the CPU gives the GPU basic directions for each frame, and then the GPU goes off to process them while the CPU prepares the next set of directions.
If the CPU isn't fast enough, the GPU sits around waiting for instructions and it causes stuttering and FPS drops. If the GPU isn't fast enough, instructions from the CPU queue up and it causes input lag. To get around the input lag, most engines intentionally drop frames and decrease FPS (basically forcing the CPU to wait) in order to keep everything in sync.
Where VRAM comes into play is that that's where textures and other important data the GPU needs are stored. The CPU and GPU can actually both write to VRAM (though the CPU does so indirectly in most cases). Writing to VRAM is an expensive operation that benefits from parallelization, especially in the context of loading textures and cached shaders, therefore more CPU cores = better performance. Meanwhile, running out of VRAM usually causes the GPU to fall back to system RAM which is much slower, so more VRAM = better performance.
TL;DR The CPU runs the game logic and the GPU renders the result, so the GPU can only render frames as fast as the CPU can prepare them and vice-versa.
Edit: Since people seem unable to tell, I'm not a bot (obligatory fuck AI); sorry I use structured formatting tho ig ¯\_(ツ)_/¯