Principal Machine Learning Engineer

Back to Jobs
Blaze Talent logo

Principal Machine Learning Engineer

Blaze Talent

200,000–250,000 / Year

Location

New York

Experience

Senior

Posted

Jul 30, 2026

Apply by

August 29, 2026

Applicants

0

Early applicantEasy applyFull-timeHybrid

Job Description

Principal ML Engineer New York, NY (Hybrid) About the Company We're building AI-native enforcement infrastructure for enterprise communication — technology that catches and fixes compliance issues in real time, before an AI-generated message ever reaches a customer or counterparty, across every channel where AI represents the business. Most existing tools only flag problems after the fact, once the risk is already out the door; we intervene before send. This is a new category, and we're the ones defining it. We're backed by top-tier venture capital and built by a team with backgrounds at major tech and financial firms, led by a founder who has built and scaled AI companies before. The Role Specialized language models sit at the core of our enforcement layer, making real-time decisions about whether a communication is safe to send. These models need to be accurate, fast, and dependable, since they operate directly in the path of live traffic. Our research team owns the underlying science — model behavior, training objectives, data strategy, and quality standards. You'll own the systems that turn that science into a reliable, production-grade product: the pipelines that train models reproducibly, the evaluation infrastructure that proves they work, and the serving stack that runs them at scale. This is a hands-on, principal-level individual contributor role on a small, senior team. It's a systems and infrastructure role, not a research role — ideal for someone who loves making ML industrial-grade. What You'll Do - Build and own training pipelines: data prep, reproducible fine-tuning runs, experiment tracking, and release automation - Build evaluation infrastructure: automated eval runs, regression gates, dashboards, and dataset versioning - Own model serving in production: low-latency inference, batching, optimization, autoscaling, and cost management - Ship model updates safely with versioning, canarying, rollback, and drift monitoring - Build repeatable workflows for adapting models to new domains and customer needs - Convert expert labels and reviewer feedback into clean training and evaluation datasets - Set the technical bar for ML infrastructure as the team grows What We're Looking For - 8+ years of software engineering experience, including 4+ years building infrastructure for ML or LLM systems in production - Hands-on depth with the modern LLM stack: PyTorch, distributed training, fine-tuning at scale (LoRA, SFT), and inference engines such as vLLM or TensorRT-LLM - Experience building eval harnesses, regression gates, or dataset pipelines, with solid understanding of precision, recall, and calibration - Proven track record owning model serving under real latency, reliability, and cost constraints — not just in notebooks - Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability - Comfort with high ownership on a small team: scoping your own work, shipping weekly, and making pragmatic build-vs-buy calls - Enjoyment of close collaboration with a research counterpart, with clear interfaces and no turf wars Nice to Have - Experience productionizing small or specialized language models - Experience with structured-output serving or constrained decoding in production - Background in a regulated or high-stakes domain such as fintech, healthcare, legal, or trust and safety - Experience deploying models into customer-controlled environments Compensation & Benefits - $200,000–$250,000 base salary, depending on experience - Performance bonus and meaningful early-stage equity - Health, dental, and vision coverage - Hybrid work from a New York office

Key Responsibilities

  • Build and own training pipelines including data prep, fine-tuning, and release automation
  • Develop evaluation infrastructure with automated runs, regression gates, and dashboards
  • Manage model serving in production with low-latency inference and autoscaling
  • Ship model updates safely using versioning, canarying, and drift monitoring
  • Create repeatable workflows for adapting models to new domains
  • Convert expert labels into clean training and evaluation datasets
  • Set technical standards for ML infrastructure

Skills Required

PythonPyTorchDistributed trainingLoRASFTvLLMTensorRT-LLMCI/CDCloud infrastructureContainersObservabilityHigh ownershipCollaborationPragmatic decision makingStructured-output servingConstrained decodingFintechHealthcareLegalTrust and safety

Benefits

  • Performance bonus
  • Equity
  • Health insurance
  • Dental insurance
  • Vision insurance

App exclusive · Free

Smart Job AI Coach

Your personal interview coach on every job — readiness tips, profile improvements, and role-specific prep. Available only in the Pulse Job app.

Interview readiness

See how prepared you are and what to improve for each role.

Personalized tips

Actionable suggestions based on your profile and the job.

After you apply

Keep coaching momentum from job detail through application success.

Get Smart Job AI Coach in the appFree on iOS and Android