Fintegrity
Back to Global Search
A

Adaption

Sourced

Inference Performance Engineer

US · San Francisco · Remote · Full-time

适配团队时区,支持远程办公

Remote worldwideOpen to China

Salary not stated

Register to see the summary, key points, full original text and apply directly

Sign up free
vLLMSGLangTensorRT-LLMCUDANCCLPythonC++RustKV-cache managementquantization

Excerpt

The role You'll own the cost and performance of our inference stack. Your work will determine how efficiently we serve models as workloads, traffic, and hardware change. You'll work closely with the engineers operating the serving fleet while owning the core performance levers: caching, batching,…

Sign up free

The summary and highlights are AI-generated and may contain errors; the original posting prevails.

Query understanding and Chinese summaries are generated by the Doubao (Yunque) large language model (generative AI service filing no. Beijing-YunQue-20230821; algorithm filing no. 网信算备110108823483901230065号), provided by Beijing Volcano Engine Technology Co., Ltd.

Sources or employers who want a posting removed or corrected can email contact@fintegrity.cn; we act within 24 hours.

Inference Performance Engineer · Adaption · Global Search | Fintegrity