All open positions
AIRemote / Hybrid 4 min read

Senior AI Engineer — LLM & Agentic Systems

Production-grade LLM systems, RAG pipelines and agentic workflows for enterprise use — end-to-end.

Employment
Full-time
Location
Warsaw / Remote (EU)
Work mode
Remote-first / Hybrid in Warsaw

Job Overview

We are building an AI Engineering team focused on production-grade LLM systems. Not another chatbot demo, not a playground that never reaches users.

Our work sits between software engineering, applied AI, data systems, cloud infrastructure, evaluation, security and product thinking. We build LLM-powered systems that run on real business data, support real users and are designed to be monitored, evaluated, secured and maintained in production.

We need a Senior or Staff-level AI Engineer who can own the full path from problem understanding to production delivery. That means architecture, model and tooling choices, implementation and shipping.

This role is a strong fit for someone who likes pragmatic engineering, understands the limits of current LLMs and knows that a reliable AI system is much more than a prompt and an API call.

What You Will Build

  • RAG systems for unstructured and semi-structured data: documents, reports, tickets, knowledge bases and CRM/ERP records.
  • Agentic workflows with tool calling, task planning, permissions, auditability and human-in-the-loop control.
  • AI assistants for operational, analytical and product teams.
  • Evaluation frameworks for answer quality, retrieval quality, prompt regressions and model behavior.
  • Integrations between LLMs, internal APIs, data warehouses and business systems.
  • Guardrails for sensitive data, hallucination control, prompt injection defense and traceability.
  • Cost, latency, model selection, fallback and routing optimization.

Responsibilities

  • Design and develop production-grade LLM systems and agentic workflows.
  • Build and optimize RAG pipelines: chunking, embeddings, retrieval, reranking, grounding and citation strategies.
  • Design agent architectures: tools, memory, planning, routing, retries, timeouts, permission models and escalation paths.
  • Integrate commercial and open-source models with applications, APIs and data systems.
  • Create evaluation frameworks: offline evaluation, regression tests, test sets, LLM-as-judge and quality gates.
  • Implement observability for AI systems: tracing, cost monitoring, latency monitoring, token usage, fallbacks and alerts.
  • Work on AI security and governance: PII redaction, data access control, prompt injection defense and audit logs.
  • Collaborate closely with Product, Data Science, Backend, Security and Cloud Engineering teams.
  • Make technology decisions: when to use LangGraph, when to build simpler orchestration, when to use RAG, classic search, fine-tuning or a hybrid approach.
  • Help shape AI Engineering standards: architectural patterns, repository templates, review checklists, documentation and delivery practices.

Technology Stack

  • Core engineering

    Python, FastAPI, Pydantic, SQL, PostgreSQL, Docker, GitHub Actions

  • LLM / AI

    OpenAI / Azure OpenAI, Anthropic, open-source LLMs, Hugging Face, embeddings, rerankers, structured outputs, function calling

  • Frameworks

    LangGraph, LangChain, LlamaIndex. Experience building lightweight custom orchestration instead of relying on frameworks by default is a strong plus

  • RAG / search

    pgvector, OpenSearch / Elasticsearch, Qdrant, Pinecone, Weaviate, hybrid search, reranking, metadata filtering

  • MLOps / LLMOps

    MLflow, Langfuse, LangSmith, OpenTelemetry, Prometheus, Grafana, test sets, golden datasets, evaluation pipelines

  • Cloud / infrastructure

    AWS, Azure or GCP, Kubernetes, Terraform, serverless, CI/CD

  • Data

    Snowflake, BigQuery, Databricks, Airflow, dbt or similar tools are a plus

Requirements

  • At least 5 years of commercial experience in software engineering, AI/ML engineering, data engineering or backend engineering.
  • Strong Python skills and solid engineering fundamentals: testing, code review, modular design, clean architecture and CI/CD.
  • Practical experience with LLM systems, RAG or agentic workflows that goes beyond local demos or notebooks.
  • Experience with LangGraph, LangChain, LlamaIndex or the ability to consciously design a simpler alternative.
  • Hands-on experience with vector databases, semantic search or hybrid search.
  • Good SQL skills and experience working with production data.
  • Ability to design systems that are reliable, observable and maintainable in production.
  • Comfortable communication in English in an international environment.
  • Ownership mindset. You care about the outcome, not just the ticket.

Nice to Have

  • Experience in regulated environments such as healthcare, fintech, insurance, legal, pharma or enterprise.
  • Experience with MLOps / LLMOps and evaluation pipelines.
  • Hands-on experience with Kubernetes, Terraform and public cloud platforms.
  • Experience with model serving tools such as vLLM, TGI, Triton or similar.
  • Understanding of LLM security risks: prompt injection, data leakage, access control and auditability.
  • Experience with large document collections, knowledge graphs or GraphRAG.
  • Previous experience as a Tech Lead, Staff Engineer or architecture owner.

What We Offer

  • Clear cooperation terms and transparent workload expectations.
  • Remote-first setup and flexible working hours.
  • Budget for conferences, training and AI/cloud certifications.
  • Equipment budget suitable for serious engineering work.
  • Access to paid AI tools, test environments and experimentation budget.
  • Time for research, prototyping and validating new approaches.
  • Private medical care and benefits package.
  • No unnecessarily long recruitment process.
  • Participation in decisions about the AI roadmap, engineering standards and team direction.

Who We're Looking For

We are looking for someone who understands that a production AI system is not just a prompt and an API call.

It is data, retrieval, evaluation, monitoring, cost, security, UX, deployment and long-term maintenance.

If you enjoy building AI systems that work beyond demo day, we should talk.

Valora

Recruitment Process

A short, focused process designed for senior practitioners.

  1. 1

    Application

    You submit your CV and consent. That's it.

  2. 2

    CV Review

    We review your background within a few business days.

  3. 3

    Technical Interview

    A focused conversation about real engineering problems.

  4. 4

    Client Interview

    Only when the role requires it, never a repeat of step 3.

  5. 5

    Offer

    Clear terms, transparent conditions, quick decision.