The mention of a NUMA node has me unclear on whether this is a software or hardware implementation. There are several vendors either offering or showing around proofs of concept to try to generate interest, for CXL memory devices(which all essentially look like a NUMA node that is nothing but memory controller); mostly aimed at hyperscaler use cases where you want the improved performance and platform features of current gen servers; but your last gen servers were beefy enough that just disposing of them would be getting rid of a relatively titanic amount of DDR4 that is still faster than most NVMe. Some of those attempt to offer software-transparent compression, some just put CXL on one side and memory controller on the other and allow you to do (mostly) agnostic mixing of memory generations on anything new enough to have some PCIe lanes support CXL extensions.
I suspect that it's mostly my ignorance that is to blame; but what I don't understand about implementing memory compression(especially the schemes that try to do it transparently with high speed general purpose compression/decompression in hardware) is how you compensate for the fact that the amount of memory you need is no longer predictable without significantly more effort(potentially untenably more if you are using RAM because latency is critical).
With normal uncompressed RAM it just demands RAM in direct proportion to its size. Potentially hideously wasteful if there's some memory-backed XML-spew log that would zip down to 10% of its nominal size; but predictable unless you've got memory safety bugs. If you are compressing your RAM suddenly you've got some uses of RAM that are functionally uncompressable and require 100% of the space their nominal size suggests they will; other things might compress exceptionally well and save you 90%; and some might change unpredictably from moment to moment depending on what needs to be stored there.
When you are doing FS compression and dedupe for backups and stuff that is normally manageable; on average enough compression is possible enough of the time that you can safely assume that you'll do better than you would if you just didn't try; and you can just keep an eye on how quickly the big SAN appears to be filling up as versions accumulate and (if tapes or other fixed-sized media are involved) just change the backup media more or less quickly; do you do that for RAM? Just use the big slack area for FS caching or something you can drop on short notice in case a burst of incompressible data comes in? Do you essentially have to rebuild how your entire workload handles RAM so that it considers, at least approximately, how well a given use of memory will compress ahead of time?
Given the sorts of files and data structures you commonly encounter I can easily enough believe that compression is worth the trouble a lot of the time; but unless you are leaving a lot of slack in your available RAM it seems like a situation where, under certain circumstances, you could go from being fine because most of what is in-memory compresses well to suddenly needing as much physical RAM as you do logical RAM because you have a bunch of already-compressed or incompressible data would markedly increase the complexity or potential for things to go wrong.