Saramin
AI Engineer (LLM / Prompt & Fine-tuning)Login to view salary
Hà Nội Chuyên viên2 years, 3 years
31 days left Be the first to apply

Department: Platform

 

About the Role

 

This role is not about wiring up an LLM API. You will own the layer that decides how our models actually behave - prompt architecture, agentic tool use, fine-tuned models, and the evaluation system that tells us whether any of it is working. If you want to spend most of your time on model behavior rather than CRUD endpoints, this is that role.

1
Your role & responsibilities

Prompt & Context Engineering (Core)

  • Design, build, and continuously iterate production prompt systems: system prompts, few-shot design, structured output enforcement.
  • Build agentic workflows: function calling / tool use, multi-step planning, error recovery, and AI harness engineering beyond simple API integration.
  • Establish prompt versioning, A/B testing, and regression test suites so prompt changes are measurable rather than guesswork.
  • Design and tune RAG pipelines: chunking strategy, embedding selection, retrieval tuning, and reranking.

 

Model Fine-tuning & Customization

  • Build training datasets (collection, cleaning, labeling, synthetic generation) and run fine-tuning experiments on open models, with team support on infrastructure.
  • Benchmark and select models across commercial APIs and open-weight models based on quality, latency, and cost.

 

Evaluation & Quality

  • Help build our evaluation practice from the ground up: golden datasets, automated scoring, LLM-as-judge
  • Define per-feature quality metrics, prevent regressions across model or prompt changes, and analyze failures such as hallucination, format breakage, and refusal.

 

Integration & Operations

  • Expose AI capabilities to product teams through clean, well-documented service interfaces.
  • Instrument observability for token cost, latency, and full request tracing.
  • Define request/response contracts for AI features with front-end and back-end engineers: streaming, failure states, and long-running jobs.
2
Your skills & qualifications

Required

  • 2+ years in software engineering, including at least 1 year building and shipping LLM-based features to production.
  • Hands-on production experience with at least one major LLM API (Anthropic, OpenAI, or Google) and at least one open-weight model family (Llama, Qwen, Mistral, or similar).
  • Hands-on fine-tuning experience (SFT, LoRA or QLoRA)
  • Strong Python, with practical use of the Hugging Face stack or PyTorch.
  • Experience designing prompts as engineered artifacts.
  • Sufficient back-end ability to ship your own work: Python or Node.js, REST APIs, and endpoints consumed by web clients (streaming, cancellation, timeout handling).

 

Preferred

  • Experience with inference optimization (vLLM, TensorRT-LLM, quantization) or GPU training infrastructure.
  • Working knowledge of vector databases and embedding-based retrieval.
  • Experience with multi-agent frameworks and orchestration.
  • Experience with Vietnamese- or Korean-language model work, or multilingual evaluation.

 

Soft Skills

  • Comfortable working with ambiguity - this field changes monthly.
  • Evidence-driven: you reach for an eval set before you reach for an opinion.
  • Positive, proactive, and highly collaborative.
  • Comfortable reading recent articles and models, and translating them into something shippable.
  • Strong analytical and problem-solving skills.

 

Application

  • CV and a link to your GitHub or Hugging Face profile.
  • A short description of one LLM feature or model you built: what the problem was, what you tried, how you measured whether it worked, and what you would do differently now. Half a page is enough.
3
Benefits
  • Gross Salary: negotiable
  • Excellent employee bonus, 13th month salary, holiday bonus, Tet...
  • 1 paid leave on Company foundation date
  • Contribute SHUI with a high percentage. Bonus with PVI health insurance.
  • Active & young working environment
  • Every 2 years of consecutive working will be added 1 more full paid annual leave
  • Annual Company trip and other team-building activities
  • Full support for AI tooling, API credits, and GPU resources required for the role

 

Working Time & Location

  • Working time: 08:30–17:30, Mon–Fri.
  • Working Location: 15th Fl, Pearl Tower, Chau Van Liem St, Tu Liem Ward, Ha Noi