I wish more database engines used a Task-based design like Umbra / CedarDB.
Most of the DB engines out there still seem to use a "n-threads" style parallelism with exchange operations and poor async I/O management.
DuckDB is improving on this front, but in some sense is catching up to R&D (and implementation!) that is now decades old.
A "rhetorical challenge" I like to give software developers working on systems like this is the following: If I gave you a computer with 1,024 cores and matching network and storage bandwidth -- but with significant latency -- could you keep a system like this 100% utilised with one query?
The answer for almost all software is "no".
For example, SQL Server tops out at 64 hardware threads for any one query: https://learn.microsoft.com/en-us/sql/database-engine/config...
GPU codes are starting to get there, but CPU codes are way behind on this frontier of computer science.
It's not just databases! Can you (de)compress a file in parallel? Verify its hash in parallel? Upload/download from storage with CPU and I/O task parallelism? Can you overlap all of these operation so nothing is ever waiting on anything else it doesn't have to?
This matters! I ran some tests with bioinformatics codes and found that most got stuck in tar pits. Many could not scale to modern SSDs with millions of IOPS or modern networking with hundreds of gigabits of throughput, no matter how many CPU cores were thrown at them.
PS: AMD's Zen 6 era EPYC 9006 processors will have 512 cores and 1,024 threads per two-socket system, so this is not hypothetical: https://www.amd.com/en/products/processors/server/epyc/9006-...
Neither one of the two databases you mentioned are open source. There is too much talk and nothing to see here.
It's impossible to parallelize crypto hash operations, if the hash covers the whole file.
You can hash chunks, or maybe use different kind of hashes though.