>with minimal lookups.
Wrong. This is one area where FPGAs have an insanely unfair advantage compared to CPUs and GPUs. Yes the SRAM is limited but you have so many individual blocks and all of them come with dual ports and getting the maximum frequency out of block RAM is much easier than getting the maximum frequency out of programmable logic.
If you wanted the highest possible memory bandwidth while being free to look up hundreds or thousands of independent memory addresses at the same time you're better off with an FPGA.
E.g. with an Efinix Titanium Ti180 you could hypothetically have 2560 simultaneous memory requests per cycle all pointing at a different address and process those requests at 1 Ghz.
Yes you are right, I didnt express what I meant very well.
looking up static values is quick and easy. When you need to lookup results from previous stages of pipeline rather than just feeding them forward, thats where I run into trouble. But I am a relatively fresh FPGA designer, so I am sure it can be done. And I probably need to level up my boards a bit too.