Jobs · smartrecruiters:inpost

I

Senior AI Engineer (m/f/n)

Inpost · Kraków, , Poland · Posted 12d ago

hybridFull-timesenior5-8 yrs🇵🇱 Poland
Apply on smartrecruiters:inpost

About the role

Job Description Why this role exists Von Halsky is InPost's conversational AI shopping assistant, live in production and serving a growing share of our customers. What decides whether it wins is not the model but whether we can tell, at release cadence, that a change made conversations better. That is the Evaluations Platform, and we are hiring the engineer who takes it to the next level. What you will own The LLM-judge pipeline. Our release-gating judge over real and golden conversations, calibrated well enough that people act on its verdicts instead of arguing about it. Eval datasets and the golden set. Real coverage across intents and categories, including the Polish-language coverage generic benchmarks do not give us. Root-cause analysis on real conversations. Making our conversation-mining stack diagnostic rather than descriptive, on a stable issue taxonomy. The "sus" detector. Abusive, adversarial and anomalous sessions, next to our guardrails and red-team work. Evals-driven development. The eval comes before the feature, and writing it is as cheap as writing the code. The interface to product and business. Vague asks in, measurable quality definitions out, and results stakeholders can act on. What "leadership aspirations" means here concretely This is a technical lead role, not a people-management role, and you will not carry line-management duties on day one. You will set and defend the technical direction for the platform, act as reviewer of record for the area, mentor other engineers, scope work with our PM and EM, present results to stakeholders, and hold the line against ad-hoc requests crowding out platform work. Engineering management later, or a Staff-level hands-on track, are both paths we will build with you. Either way we need someone accountable for an area rather than for a ticket. How we define success in this role Judge pass rate is calibrated against human labels and is a number people trust and cite. A stable, versioned issue taxonomy is live and week-over-week trends are comparable. Every release is gated by an eval run the team can reproduce. The golden set has documented coverage and named blind spots. At least two engineers besides you can operate and extend the platform. Qualifications What we are looking for Required 5+ years building and running production software, with 2+ years on LLM-based systems that real users hit. Strong engineering fundamentals , plus the habits that go with production ownership: testing, CI/CD, containers, observability, and working in cloud. We work primarily in Python. Demonstrable experience evaluating generative systems, not only building them: LLM-as-judge, human-label calibration, inter-annotator agreement, regression suites, offline-versus-online divergence. You should have opinions about what makes an eval worthless. AI engineering fundamentals. Prompt and context engineering as a discipline (context-window budgeting, structured outputs, failure-mode taxonomies); agentic primitives in production (tool use, multi-turn state, MCP , agent-to-agent integration patterns); and eval and LLM-observability tooling (LangFuse, Braintrust, Weave or equivalent, including things you built yourself). Comfort with data at scale : SQL, working with a lake or warehouse, and building a metric someone else can reproduce. Fluency with AI-assisted development tooling (Claude Code, Cursor, Copilot). We use it daily and expect it. Ability to make a technical argument to a non-technical audience and be understood. English B2 and Polish . Our users converse in Polish and you will read their conversations; judging quality you cannot read is not possible. Nice to have Harness and loop engineering. Building the scaffolding around models rather than only calling them: agent loops, retries and fallbacks, tool-call orchestration, deterministic replay, and the plumbing that makes a non-deterministic system testable. Auto-improving systems. Closing the loop from production signal back into the product: mining failures into cases, using eval results to drive prompt, retrieval and routing changes, and automating the parts of that cycle that people do by hand today. Adversarial robustness, jailbreak testing, red-teaming, or abuse and fraud detection. E-commerce, search or recommendation domain experience. Company Description InPost Group is an innovative European out of home deliveries company, revolutionizing the way parcels are delivered to customers. With operations across several countries, our network of intelligent lockers provides customers with a fast, convenient, and secure delivery option. InPost Group is a publicly traded company, with a market capitalization of about $5 billion as of March 2023. With over 10,000 people worldwide, InPost Group is one of the largest out of home delivery providers in Europe, committed to providing sustainable and efficient delivery solutions to meet the evolving needs of customers in today's rapidly changing landscape. We are a team that

Read the full posting on smartrecruiters:inpost →

FAQ

Is the Senior AI Engineer (m/f/n) role at Inpost remote?+

This Senior AI Engineer (m/f/n) position is listed as hybrid (Kraków, , Poland).

What seniority level is this Senior AI Engineer (m/f/n) role?+

This is a senior level position.

How do I apply for the Senior AI Engineer (m/f/n) role at Inpost?+

Use the "Apply on smartrecruiters:inpost" button to open the original posting on smartrecruiters:inpost, where you can submit your application directly to Inpost.