Fintegrity
Back to Global Search
P

Plaud

Reliable source

Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco

US · San Francisco · San Francisco, CA · Remote · Full-time

需适配美国旧金山时区

Remote worldwideAI & MLOpen to China

Salary not stated

Register to see the summary, key points, full original text and apply directly

Sign up free
LLM inference enginespeech modelKV cache managementPagedAttentionGPU architecturesNVIDIA AmpereNVIDIA HoppervLLMTensorRT-LLMSGLangNVIDIA Triton Inference ServerWebSocketsWebRTCspeculative decodinglookahead decodingmodel compressionquantizationFP8INT8AWQ

Excerpt

About Plaud Inc. Plaud is building the real-world AI interface for professionals to amplify intelligence, elevate productivity and performance, loved by over 2,500,000 users worldwide since 2023. With a mission to amplify human intelligence, Plaud captures, structures, and compounds the…

Sign up free

The summary and highlights are AI-generated and may contain errors; the original posting prevails.

Query understanding and Chinese summaries are generated by the Doubao (Yunque) large language model (generative AI service filing no. Beijing-YunQue-20230821; algorithm filing no. 网信算备110108823483901230065号), provided by Beijing Volcano Engine Technology Co., Ltd.

Sources or employers who want a posting removed or corrected can email contact@fintegrity.cn; we act within 24 hours.

Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco · Plaud · Global Search | Fintegrity