Formal
Senior AI Engineer · Aleph AlphaDear Hiring Manager,
Operating reliable, high-throughput inference for large language models while keeping cost and latency predictable is the practical problem Aleph Alpha faces when delivering enterprise AI on European infrastructure. At Delivery Hero I led the Partner Support Triage Assistant, a retrieval-augmented system over 1.4 million historical tickets that cut median time-to-owner from 4 hours 20 minutes to 1 hour 35 minutes and reduced misrouted tickets by two-thirds, while halving inference cost through speculative decoding and dynamic batch routing. That project taught me where production risk concentrates and how to measure it before and after changes.
Most of my time there focused on making the inference and retrieval stack dependable under load. I designed and shipped a hybrid retrieval pipeline combining BM25, pgvector, and cross-encoder re-ranking, and operationalised it with an evaluation playbook that runs offline benchmarks, regression checks, and safety probes before rollout. I also integrated speculative decoding across self-hosted vLLM instances and a hosted API to control tail cost, and implemented dynamic batching to smooth p95 latency. The end-to-end work included FastAPI-based serving, PostgreSQL with pgvector for the vector store, and Kubernetes for orchestration, and it became the team standard for on-call runbooks and release gates.
Sincerely, Alex Morgan
