Site Reliability Engineer (SRE) On-Prem
Back to Jobs
DreamGet Smart Job AI Coach in the appFree on iOS and Android 






Site Reliability Engineer (SRE) On-Prem
Location
TLV - ISR
Experience
Mid
Posted
Jul 22, 2026
Apply by
August 21, 2026
Applicants
0
Early applicantEasy applyFull-timeWork from Office
Job Description
Every nation has data. Few can protect it. Fewer still can act on it.
Dream is the sovereign AI and national cyber-defense company for governments.
We help nations secure their most critical systems, connect fragmented information at a national scale, and turn their most sensitive data into decisions, all fully sovereign.
This is more than a job. It's a Dream job, where you'll work at a global scale alongside some of the best AI researchers, cyber operators, and government experts in the world.
The mission only works if the company behind it does. This role keeps Dream running at the scale our work demands. And our work demands a uniquely global scale.
We are on an expedition to find an On-Premise Site Reliability Engineer (SRE)- someone who is passionate about building rock-solid, high-performance infrastructure and bringing order to complex environments. In this role, you will own the end-to-end reliability, automation, and deployment of DREAM’s platform across customer sites, working hands-on with cutting-edge AI, bare-metal, and hybrid cloud architectures.
You’ll collaborate closely with Product, R&D, and Architecture teams while serving as the ultimate technical authority for our customer deployments. From designing automated Ansible workflows and mastering Kubernetes to troubleshooting complex network topologies, you will eliminate toil, streamline cluster operations, and ensure every deployment is scalable, seamless, and mission-ready.
- Lead End-to-End On-Prem & Hybrid Deployments: Own the technical delivery and reliability of DREAM’s platform in close collaboration with Product, R&D, and customer technical teams.
- Architect, Execute & Improve K8s Deployments: Take a definitive hands-on role in deploying, configuring, operating, and continuously improving our platform using advanced, enterprise-grade Kubernetes architectures.
- Helm Chart Management: Design, modify, and manage Helm charts to package, version, and streamline complex application deployments across different environments.
- Drive Automation & Simplification: Design, implement, and maintain robust deployment automation using Ansible. You must have a passion for turning complex manual tasks into reliable, repeatable, single-click operations.
- Manage Infrastructure as Code: Utilize Git as the single source of truth to manage configurations, manifests, and automation playbooks, enforcing modern engineering best practices.
- Bridge On-Prem and Cloud: Leverage AWS resources (specifically EC2 and S3) for hybrid components, staging environments, or cloud-to-on-prem data flows.
- Technical Tier-3 Escalation: Serve as the ultimate technical authority for deployment, Linux networking, and Kubernetes orchestration issues.
- Continuous Improvement: Constantly refine our delivery pipelines, optimize bootstrap processes, and create rock-solid technical documentation.
- SRE / Delivery Mindset: 3–5 years of hands-on experience in enterprise infrastructure deployment, systems engineering, or an on-prem operational reliability role.
- Kubernetes & Helm Expert: Deep, production-grade experience with Kubernetes architecture, deployment, advanced troubleshooting, and CNI networking. Proven working experience creating, maintaining, and deploying applications using Helm charts.
- Ansible Mastery: Proven experience writing clean, scalable Ansible roles and playbooks for configuration management, automation, and infrastructure provisioning.
- Modern Workflows (Git & AWS): Solid working experience using Git for version control and collaborating on code/infrastructure. Practical experience provisioning and managing AWS resources (EC2 and S3).
- Core Systems & Linux: Strong Linux background (Ubuntu) with a deep understanding of system internals, containerized runtimes, and troubleshooting distributed applications.
- Solid Networking Knowledge: Hands-on experience with routing, firewalls, and switching topology (mainly Cisco)
- Storage Foundations: Working knowledge of storage protocols (iSCSI, SAN, local NVMe) and enterprise storage arrays (like DELL) interacting with Kubernetes Persistent Volumes.
- GPU & Accelerated Compute: Working knowledge of managing GPU-enabled Kubernetes nodes, including NVIDIA drivers/runtime and basic troubleshooting.
- Air-Gapped Deployments: Experience deploying and maintaining software in air-gapped or offline environments, including registry mirroring and artifact staging.
- Problem-Solver: Strong debugging and problem-solving skills in complex, distributed environments with an intense ownership and accountability mindset.
- Willingness to Travel: Ready to travel to customer sites for physical staging and deployments—at least 30%.
- Language: High-level English proficiency.
Nice to have
- EU or any other additional citizenship.
- Experience with enterprise data components like MongoDB, PostgreSQL, Neo4j, or RabbitMQ.
- Working experience with project management tools such as Jira and Monday.
- Valid Israeli security clearance.
If you think this role doesn't fully match your skills but are eager to grow and break glass ceilings, we’d love to hear from you!
Key Responsibilities
- Own technical delivery and reliability of the platform in close collaboration with Product, R&D, and customer teams.
- Deploy, configure, operate, and improve platform using advanced Kubernetes architectures.
- Design, modify, and manage Helm charts for application deployments.
- Implement and maintain robust deployment automation using Ansible.
- Manage infrastructure as code using Git as the single source of truth.
- Leverage AWS resources for hybrid components and cloud-to-on-prem data flows.
- Serve as technical authority for deployment, Linux networking, and Kubernetes orchestration issues.
- Refine delivery pipelines, optimize bootstrap processes, and create technical documentation.
Skills Required
KubernetesHelmAnsibleGitAWSEC2S3LinuxUbuntuNetworkingCiscoiSCSISANNVMeGPUNVIDIAAir-gapped deploymentsProblem solvingOwnershipAccountabilityCommunicationMongoDBPostgreSQLNeo4jRabbitMQJiraMonday
App exclusive · Free
Smart Job AI Coach
Your personal interview coach on every job — readiness tips, profile improvements, and role-specific prep. Available only in the Pulse Job app.
Interview readiness
See how prepared you are and what to improve for each role.
Personalized tips
Actionable suggestions based on your profile and the job.
After you apply
Keep coaching momentum from job detail through application success.
Similar roles for you
Matched using this role's title and skills. Open the job search anytime to see every listing.

NEW- Applied AI Engineer, Site Reliability Engineer - EMEA
Mistral
Full-timeEasy applyWork from Home
ParisSenior

IT Operations Engineer
Datamatics Technologies
Full-timeEasy applyWork from Office
RiyadhMid

IT Operations Engineer
Datamatics Technologies
Full-timeEasy applyWork from Office
BengaluruMid

Contract Lead, Site Reliability Engineering — AI Accelerator Infrastructure
d-Matrix
195,000–285,000 / Year
ContractEasy applyWork from Home
Santa ClaraSenior

IT Operations Engineer
Datamatics Technologies
Full-timeEasy applyWork from Office
KarachiSenior

IT Operations Engineer
Datamatics Technologies
Full-timeEasy applyWork from Office
CairoSenior