Mingxin Technology
Storage acceleration for LLM inference. KV-cache tiering on NVMe-oF: +29-40% throughput (signed benchmarks, reproducible). mingxinstorage.xyz
loading...
Storage acceleration for LLM inference. KV-cache tiering on NVMe-oF: +29-40% throughput (signed benchmarks, reproducible). mingxinstorage.xyz