Hiring

AI Inference Platform — Bailian Model Serving Runtime

Job Description

Design and build the next-generation distributed AI inference runtime: make large-model services start faster, run more reliably, and use resources more efficiently. You will work on the core engineering problems of large-scale inference systems, including:

  1. Ultra-fast elasticity of inference instances: continuously cut model loading and instance startup time through container image acceleration, model and data caching, P2P distribution, prefetching, and checkpoint/restore, improving the elastic scaling efficiency of large clusters;
  2. Heterogeneous clusters and distributed inference: run different models and inference engines on GPUs, NPUs, and other accelerators; optimize computation, communication, and state coordination across nodes; explore more efficient parallel and distributed inference architectures;
  3. Day-0 model launch: work closely with model, engine, and hardware teams to understand the system requirements brought by new model architectures, and own the full path from runtime adaptation and performance tuning to production-scale deployment, so that new models go live quickly and reliably;
  4. Performance, stability, and cost optimization: continuously improve throughput, latency, availability, and resource utilization through system architecture optimization, performance analysis, and engineering innovation, keeping large-scale AI applications running reliably.

Requirements

  1. Graduate degree or above in computer science or a related field, with solid CS fundamentals;
  2. Strong system design and engineering skills; proficient in at least one mainstream language (C++ / Go / Java / Python / Rust), able to design and implement complex system modules;
  3. Familiar with Linux system development environments, with good engineering practice: system tuning, performance analysis, and problem diagnosis;
  4. Working understanding of common inference engines such as vLLM, SGLang, and llama.cpp;
  5. Strong interest in AI infrastructure or large-model inference systems, with good learning ability and curiosity to explore new techniques;
  6. Bonus: papers at top systems conferences, international ACM/ICPC awards, or active contributions to major open-source projects.

Contact Us

Campus hires (top candidates may apply for Alibaba Star)

Experienced hires

Comment on GitHub →

Built with Hugo
Theme Stack designed by Jimmy