Skip to content
View iemAnshuman's full-sized avatar

Highlights

  • Pro

Block or report iemAnshuman

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
iemAnshuman/README.md

Anshuman Agrawal

Distributed AI Systems Researcher
Accelerator communication, memory movement, and collective operations.

Google Summer of Code 2026 20+ merged HPX pull requests A Square Blog


About

I work on distributed AI data-plane systems: the software that moves tensors and model state across accelerator memory, GPUs, nodes, and high-speed networks.

My interests include collective communication, GPU-aware transfers, topology-aware transport, communication–computation overlap, and reproducible performance analysis.

I am a Google Summer of Code 2026 contributor with the STE||AR Group, working on hpx::collectives.


Selected work

  • Implemented hierarchical all_reduce, all_gather, all_to_all, and prefix scans.
  • Diagnosed serialization, centralized data-path, and transport-threshold bottlenecks.
  • Reduced large-message all_to_all performance from 7.1× behind OpenMPI to approximately 1.2×.
  • Added contiguous multidimensional payloads, communicator-generation management, benchmarks, and distributed regression tests.

A model-free regression canary for distributed-LLM communication that preserves configuration rankings, regression decisions, and latency-tail behaviour.

communication trace → canary → replay → verify

Current focus

  • Accelerator data movement and GPU-aware communication
  • Collective algorithms and distributed runtime systems
  • Memory registration, staging, and asynchronous transfers
  • Communication–computation overlap
  • Cluster-scale performance profiling

Stack

C++20 · Python · CUDA · HPX · MPI · NCCL · LCI · Triton · Linux · Slurm


Email · Blog · X · LinkedIn

Popular repositories Loading

  1. neuro-ranker-distill neuro-ranker-distill Public

    Python 7 2

  2. Data_Structures Data_Structures Public

    C

  3. HMS-python HMS-python Public

    Python

  4. Research-Internships-for-Undergraduates Research-Internships-for-Undergraduates Public

    Forked from zapplyjobs/Research-Internships-for-Undergraduates

    List of Research Internships for Undergraduate Students

  5. Todoist-CLI Todoist-CLI Public

    Python

  6. oop-uni-codes oop-uni-codes Public

    Java