Principal Engineer, AI Serving Framework Architect (Software)

samsungsemiconductor · San Jose, California, United States

Posted
23 days ago
Last confirmed live
2 days ago

What this role involves

This Principal Engineer role focuses on architecting AI serving frameworks for large-scale systems. The position involves leading research on dynamic scheduling, memory optimization, and inference acceleration for multi-rack scale systems. Required expertise includes LLM inference stacks, vLLM, and proficiency in PyTorch, Python, and C++.

Skills this posting asks for

  • ai serving framework
  • large language model inference
  • vllm
  • pytorch
  • python
  • c++
  • compute-capable memory
  • hierarchical memory
  • rag
  • vector db
  • ai agent
  • knowledge-graph
  • kvcache
  • optimization algorithms
  • multi-rack scale systems
  • heterogeneous devices
  • profiling
  • dynamic scheduling
  • memory optimization
  • inference acceleration

Requirements

  • 10 years of experience
  • Level: principal
  • Remote policy: onsite

From the employer’s posting

Please Note: To provide the best candidate experience amidst our high application volumes, each candidate is limited to 10 applications across all open jobs within a 6-month period.&nbsp; <span style=…

Read the full description on samsungsemiconductor’s careers page

Apply without filling the form

Approve this role and the application is completed for you, including a résumé tailored to it. You get a confirmation when it lands, and a credit is only spent when a submission is confirmed.

Other roles at samsungsemiconductor

All 10 roles at samsungsemiconductor