Skip to content

Infera Development Roadmap (2026 Q3) #9

Description

@jiejingzhangamd

Infera Development Roadmap (2026 Q3)

Planned work by category — current state and the target.

Category Item Today Goal
Model support Multimodal (image / audio / video) Routed on load only (no cache locality) MM-aware routing; encode/prefill/decode (EPD) separation
Hardware support MI325 / MI300 Validated on MI355X only Validate + CI on MI325 / MI300
PD / networking / comms Transports RDMA (Mooncake / MoRI / NIXL); xGMI diagnostics only Evaluate xGMI / remote-copy for intra-node KV
Parallelism Wide expert parallelism (EP) Not supported 32+ GPU instances spanning nodes
KV-cache management Offload engine coverage vLLM only (incl. AIC GPU-Direct) SGLang (incl. GPU-Direct), then ATOM
KV-cache management Cluster-wide KV pool L3 or L4 (either/or) Composed L3 + L4; multi-model namespacing
KV-cache management Distributed prefill cache Per worker Shared across prefill workers and with decoders
Operator optimization Op injection HyperLoom Optimized OP Injection Cross-engine op/kernel injection using HyperLoom optimized by new model and user workload
Framework Dynamic scaling (Kubernetes) Static fleet; workers self-register, no autoscaler Load-driven autoscaling; runtime role switching
Framework SLA-aware scheduling Relative cost heuristic; throughput only Settable SLO targets and goodput reporting

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions