LLM Engineer (Reinforcement Learning)
Back to Jobs
42dotGet Smart Job AI Coach in the appFree on iOS and Android 

LLM Engineer (Reinforcement Learning)
Location
Pangyo (Software Dream Center), South Korea
Experience
Mid
Posted
Jul 22, 2026
Apply by
August 21, 2026
Applicants
0
Early applicantEasy applyFull-timeWork from Home
Job Description
# About the Team & Mission
LLM Engineer(Reinforcement Learning)는 LLM학습 파이프라인을 설계하여 실서비스에서 활용 가능한 생성형 언어모델을 학습합니다.
지속적인 품질 향상을 위하여 끊임없이 새로운 방법론을 시도하여, 실사용자에게 꼭 필요한 서비스를 출시하고, LLM 스스로 품질을 개선할 수 있도록 가다듬는 일에 기여합니다.
## Responsibilities
- LLM학습 과정의 효율 향상
- PLM 또는 Fine-tuned LLM의 Direct Alignment Algorithm / PPO, GRPO, DPO 등을 이용한 학습 과정의 전반적인 효율 향상
- 생성 결과의 전반적인 정확성과 안정성 향상
- 생성 결과의 품질 향상을 위하여 Reward Hacking을 방지하고, Self-Refine이 가능한 학습 구조 설계
- 외부 지식 및 API와 연동 가능한 기초 모델 개발
- 지시의 종류에 따라 스스로 필요한 외부 연동 Tool을 선택하는 LLM 학습
## Qualifications
- Deep Learning 또는 NLP 관련 경력 3년 이상 (석사 신입 지원 가능)
- 숙련된 프로그래밍 (Python & pytorch) 능력
- PyTorch를 활용한 모델 설계, 학습, 평가 및 최적화 경험
- GPU를 활용한 LLM 학습 및 Trouble shooting 능력
- 분산 학습 프레임워크(Slurm, DDP, Horovod 등) 사용 경험
- 동료와의 원활한 협업 능력
## Preferred Qualifications
- Deep Learning/NLP 관련 논문 제출 또는 석박사 학위 소지자
- 주요 학술 대회(ACL, EMNLP, NeurIPS 등) 논문 발표 경험
- Docker 및 Kubernetes에 대한 경험
- GPU 클러스터를 활용한 학습 파이프라인 설계 및 관리 경험
- GPU를 활용한 학습 및 서비스 개발 경험
- GPU 기반의 Training 또는 Inference 시스템 구축 경험
- LLM의 Post-training 관련 경험
- Supervised Fine-Tuning 및 Parameter Efficient Fine-Tuning 활용 경험
## Interview Process
- 서류 전형
- 코딩 테스트
- 1차 면접 (화상, 1시간 내외)
- 2차 면접 (대면 혹은 화상, 3시간 내외)
- 처우 협의·입사
## Additional Information
- 전형 절차는 일정 및 진행 상황에 따라 일부 변경될 수 있으며, 각 전형 결과는 등록하신 이메일로 개별 안내드립니다.
- 지원서 제출 시 주민등록번호, 가족관계, 혼인 여부, 연봉, 사진, 신체조건, 출신 지역 등 채용절차법상 요구 금지된 정보는 제외 부탁드립니다.
- 지원서 접수 중 오류가 발생하거나 기타 문의 사항이 있을 경우, recruit@42dot.ai로 문의해 주시기 바랍니다.
- 국가보훈대상자 및 취업보호 대상자는 관계법령에 따라 우대합니다.
- 장애인 고용 촉진 및 직업재활법에 따라 장애인 등록증 소지자를 우대합니다.
- 42dot은 의뢰하지 않은 서치펌의 이력서를 받지 않으며, 요청하지 않은 이력서에 대해 수수료를 지불하지 않습니다.
- 지원서 내용 중 허위 사실이 발견될 경우, 입사가 취소될 수 있습니다.
- 인터뷰 프로세스 종료 후 지원자의 동의하에 평판조회가 진행될 수 있습니다.
- 3개월의 수습기간이 적용될 수 있습니다.
Key Responsibilities
- Improve efficiency of LLM learning processes using Direct Alignment Algorithms, PPO, GRPO, and DPO.
- Enhance accuracy and stability of generation results by preventing reward hacking and designing self-refine learning structures.
- Develop foundational models capable of integrating with external knowledge and APIs.
Skills Required
PythonPyTorchDeep LearningNLPGPUSlurmDDPHorovodCollaborationDockerKubernetesGPU cluster managementTraining system developmentInference system developmentSupervised Fine-TuningParameter Efficient Fine-Tuning
App exclusive · Free
Smart Job AI Coach
Your personal interview coach on every job — readiness tips, profile improvements, and role-specific prep. Available only in the Pulse Job app.
Interview readiness
See how prepared you are and what to improve for each role.
Personalized tips
Actionable suggestions based on your profile and the job.
After you apply
Keep coaching momentum from job detail through application success.
Similar roles for you
Matched using this role's title and skills. Open the job search anytime to see every listing.

Data Scientist - Agentic AI Systems - IFS Loops
IFS
140,000–150,000 / Year
Full-timeEasy applyHybrid
Palo AltoMid
Software Engineer, GenAI Silicon Automation, DeepMind
Google
82.000–84.500 / Year
Full-timeEasy applyWork from Office
ParisMid
Staff Software Engineer, AI-Powered GRC Automation
Google
207,000–301,000 / Year
Full-timeEasy applyWork from Office
SunnyvaleSenior
Machine Learning Engineer
Trilyon, Inc.
$60–$85 / Hour
ContractEasy applyWork from Home
Remote
Senior Software Engineer, Machine Learning, Acceleration Platform
Google
Full-timeEasy applyWork from Office
SingaporeSenior
Software Engineer III, AI/ML Shopping Creator, YouTube
Google
Full-timeEasy applyWork from Office
ZürichMid