Jobs · empllo

T

LLM Inference Frameworks and Optimization Engineer

Aperture Cloud · San Francisco · Posted 3d ago

remotemid181000-260000 USD
Apply on empllo

Available in 2 locations

San Francisco · remote Apply → San Francisco, Singapore, Amsterdam · onsite Apply →

About the role

📋 Description Design fault-tolerant, high-concurrency inference engine for text, image, and multimodal models. Implement distributed inference strategies: MoE parallelism, tensor parallelism, and pipeline parallelism. Apply CUDA graph, TensorRT graph optimizations, and torch.compile for speed. Collaborate with hardware teams to identify bottlenecks and co-optimize GPU inference. Work with AI researchers and infra engineers to optimize end-to-end model serving. 🎯 Requirements 3+ years in deep learning inference, distributed systems, or HPC. Familiar with at least one LLM inference framework (TensorRT-LLM, vLLM, SGLang, TGI). Background in GPU programming (CUDA/Triton/TensorRT), compilers, quantization, and GPU cluster scheduling. Deep understanding of KV cache systems like Mooncake, PagedAttention, or in-house variants. Proficient in Python and C++/CUDA for high-performance DL inference. Transformer/LLM/VLM/Diffusion model optimization; workload scheduling, CUDA graph, compiled kernels. 🎁 Benefits Experience with RDMA/RoCE in large-scale data center networks. Familiar with distributed file systems (3FS, HDFS, Ceph). Familiar with Kubernetes and open-source distributed scheduling/orchestration. Contributions to open-source deep learning inference projects.

Read the full posting on empllo

FAQ

Is the LLM Inference Frameworks and Optimization Engineer role at Aperture Cloud remote?+

This LLM Inference Frameworks and Optimization Engineer position is listed as remote (San Francisco).

What is the salary for the LLM Inference Frameworks and Optimization Engineer role at Aperture Cloud?+

The listing states 181000-260000 USD.

What seniority level is this LLM Inference Frameworks and Optimization Engineer role?+

This is a mid level position.

How do I apply for the LLM Inference Frameworks and Optimization Engineer role at Aperture Cloud?+

Use the "Apply on empllo" button to open the original posting on empllo, where you can submit your application directly to Aperture Cloud.