Model Serving Engineer
Back to Jobs
BrightvisionGet Smart Job AI Coach in the appFree on iOS and Android

Model Serving Engineer
74,000–98,000 / Year
Location
Remote
Experience
Mid
Posted
Jul 30, 2026
Apply by
August 29, 2026
Applicants
0
Early applicantEasy applyFull-timeWork from Home
Job Description
Model Serving Engineer – Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Model Serving Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $74,000–$98,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are seeking a Model Serving Engineer to design, build, and operate high-performance, highly reliable inference platforms for serving large machine learning models in production. The role focuses on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse model workloads. The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale, and understands the trade-offs between latency, throughput, cost, and quality in ML serving.
Required Qualifications
- Bachelor’s or Master’s degree in Computer Science or a related field.
- Six or more years of experience in distributed systems, infrastructure, or ML platform engineering.
- Strong proficiency in Python and a systems language such as Go, Rust, or C++.
- Deep experience operating high-throughput, low-latency services in production.
- Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM.
- Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
- Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
- Experience with observability stacks including metrics, tracing, and structured logging.
- Solid grounding in performance engineering and capacity planning.
- Strong communication and incident response skills.
Preferred Qualifications
- Open-source contributions to model serving infrastructure.
- Experience with multi-region or globally distributed AI serving.
- Familiarity with model quantization, distillation, and compression techniques.
- Exposure to FinOps for AI workloads and cost-efficient serving design.
- Experience supporting external-facing AI APIs at scale.
How to Apply
Would you like to know more about this opportunity?
For immediate consideration, please send your resume to [\[email protected\]](/cdn-cgi/l/email-protection)
Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.
Equal Employment Opportunity (EEO) Statement
Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.
BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.
Key Responsibilities
- Design and build high-performance inference platforms for serving large machine learning models.
- Implement request routing, batching, caching, autoscaling, and GPU utilization strategies.
- Ensure end-to-end observability across diverse model workloads.
- Perform performance engineering and capacity planning for ML serving systems.
- Manage incident response and maintain system reliability in production environments.
Requirements
- Bachelor's or Master's degree in Computer Science or a related field
Skills Required
PythonGoRustC++vLLMTensorRT-LLMKubernetesGPU architectureMetricsTracingStructured loggingPerformance engineeringCapacity planningCommunicationIncident responseOpen-source contributionsMulti-region AI servingModel quantizationModel distillationModel compressionFinOpsAI API support
App exclusive · Free
Smart Job AI Coach
Your personal interview coach on every job — readiness tips, profile improvements, and role-specific prep. Available only in the Pulse Job app.
Interview readiness
See how prepared you are and what to improve for each role.
Personalized tips
Actionable suggestions based on your profile and the job.
After you apply
Keep coaching momentum from job detail through application success.

