Model Serving Engineer

Back to Jobs
Brightvision logo

Model Serving Engineer

Brightvision

74,000–98,000 / Year

Location

Remote

Experience

Mid

Posted

Jul 30, 2026

Apply by

August 29, 2026

Applicants

0

Early applicantEasy applyFull-timeWork from Home

Job Description

Model Serving Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. Job Title: Model Serving Engineer Location: 100% Remote (U.S.) Position Type: Full-time, Direct W2 Salary Range: $74,000–$98,000 Annually Experience Required: 6+ years Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position. Job Summary We are seeking a Model Serving Engineer to design, build, and operate high-performance, highly reliable inference platforms for serving large machine learning models in production. The role focuses on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse model workloads. The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale, and understands the trade-offs between latency, throughput, cost, and quality in ML serving. Required Qualifications - Bachelor’s or Master’s degree in Computer Science or a related field. - Six or more years of experience in distributed systems, infrastructure, or ML platform engineering. - Strong proficiency in Python and a systems language such as Go, Rust, or C++. - Deep experience operating high-throughput, low-latency services in production. - Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM. - Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization. - Familiarity with Kubernetes, autoscaling, and modern cloud platforms. - Experience with observability stacks including metrics, tracing, and structured logging. - Solid grounding in performance engineering and capacity planning. - Strong communication and incident response skills. Preferred Qualifications - Open-source contributions to model serving infrastructure. - Experience with multi-region or globally distributed AI serving. - Familiarity with model quantization, distillation, and compression techniques. - Exposure to FinOps for AI workloads and cost-efficient serving design. - Experience supporting external-facing AI APIs at scale. How to Apply Would you like to know more about this opportunity? For immediate consideration, please send your resume to [\[email protected\]](/cdn-cgi/l/email-protection) Learn more about Bright Vision Technologies at www.bvteck.com. Bright Vision Technologies is an Equal Opportunity Employer. Equal Employment Opportunity (EEO) Statement Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall. BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.

Key Responsibilities

  • Design and build high-performance inference platforms for serving large machine learning models.
  • Implement request routing, batching, caching, autoscaling, and GPU utilization strategies.
  • Ensure end-to-end observability across diverse model workloads.
  • Perform performance engineering and capacity planning for ML serving systems.
  • Manage incident response and maintain system reliability in production environments.

Requirements

  • Bachelor's or Master's degree in Computer Science or a related field

Skills Required

PythonGoRustC++vLLMTensorRT-LLMKubernetesGPU architectureMetricsTracingStructured loggingPerformance engineeringCapacity planningCommunicationIncident responseOpen-source contributionsMulti-region AI servingModel quantizationModel distillationModel compressionFinOpsAI API support

App exclusive · Free

Smart Job AI Coach

Your personal interview coach on every job — readiness tips, profile improvements, and role-specific prep. Available only in the Pulse Job app.

Interview readiness

See how prepared you are and what to improve for each role.

Personalized tips

Actionable suggestions based on your profile and the job.

After you apply

Keep coaching momentum from job detail through application success.

Get Smart Job AI Coach in the appFree on iOS and Android