🍊
GPU kernel engineer @ Z.ai. LLM inference & training performance
Pinned Loading
-
sgl-project/sglang
sgl-project/sglang PublicSGLang is a high-performance serving framework for large language models and multimodal models.
-
NVIDIA/TensorRT-LLM
NVIDIA/TensorRT-LLM PublicTensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…
-
NVIDIA/cutlass
NVIDIA/cutlass PublicCUDA Templates and Python DSLs for High-Performance Linear Algebra
-
cudnn-frontend
cudnn-frontend PublicForked from NVIDIA/cudnn-frontend
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.

