Infera Development Roadmap (2026 Q3)
Planned work by category — current state and the target.
| Category |
Item |
Today |
Goal |
| Model support |
Multimodal (image / audio / video) |
Routed on load only (no cache locality) |
MM-aware routing; encode/prefill/decode (EPD) separation |
| Hardware support |
MI325 / MI300 |
Validated on MI355X only |
Validate + CI on MI325 / MI300 |
| PD / networking / comms |
Transports |
RDMA (Mooncake / MoRI / NIXL); xGMI diagnostics only |
Evaluate xGMI / remote-copy for intra-node KV |
| Parallelism |
Wide expert parallelism (EP) |
Not supported |
32+ GPU instances spanning nodes |
| KV-cache management |
Offload engine coverage |
vLLM only (incl. AIC GPU-Direct) |
SGLang (incl. GPU-Direct), then ATOM |
| KV-cache management |
Cluster-wide KV pool |
L3 or L4 (either/or) |
Composed L3 + L4; multi-model namespacing |
| KV-cache management |
Distributed prefill cache |
Per worker |
Shared across prefill workers and with decoders |
| Operator optimization |
Op injection |
HyperLoom Optimized OP Injection |
Cross-engine op/kernel injection using HyperLoom optimized by new model and user workload |
| Framework |
Dynamic scaling (Kubernetes) |
Static fleet; workers self-register, no autoscaler |
Load-driven autoscaling; runtime role switching |
| Framework |
SLA-aware scheduling |
Relative cost heuristic; throughput only |
Settable SLO targets and goodput reporting |
Infera Development Roadmap (2026 Q3)
Planned work by category — current state and the target.