Ibin Mathew
Ibin Mathew Biju standing on a snow-covered path in an Alpine village
Open to new roles & collaborationsMannheim, Germany

Ibin Mathew
Biju

I build systems that think, scale, and ship — across AI, data, and software.

(01)What I do

Focus areas

Where I go deep — from prototype to a system you can measure, monitor and trust.

Generative AI

LLM applications, RAG pipelines, agentic and multi-agent systems (LangGraph, LangChain, MCP), prompt engineering, evaluation and guardrails. Built and industrialised a production agent system end to end.

LangGraphLangChainRAGMCPMulti-agent

Natural Language Processing

Document extraction and structuring from unstructured sources, retrieval over large text corpora, hybrid search, and text classification models.

Document AIHybrid searchEmbeddingsTransformers

MLOps

Evaluation and monitoring pipelines (Langfuse), automated validation, benchmarking and regression detection, containerised deployment with Docker and Kubernetes, and CI/CD.

LangfuseDockerKubernetesCI/CDBenchmarking
(02)Currently

What I'm building now

The German token tax

Case study

A short research build measuring the hidden cost of running LLMs in German: I tokenized 24 English/German prompt pairs across 9 tokenizer families and mapped where the overhead compounds. Same meaning, far more tokens — and since models bill per token, a real bill.

+68% more tokens for Germanup to 2× on Claude9 tokenizers · 24 prompts
Read the case study
(03)Selected work

Featured projects

Case study

The German token tax

How much more do LLMs cost in German than English? I measured token overhead across 9 tokenizer families and 24 prompt pairs — German needs +68% more tokens for the same meaning, and up to 2× on Claude.

LLMsTokenizersData vizAnalysis
View case study
Patent · Filed, pending

Multi-Agent Orchestration with Memory and Validation

Co-author on a filed patent covering the orchestration of multiple AI agents with shared memory and built-in validation — the architecture behind the production agent system I built at NEC.

Multi-agentLangGraphPatent

PubMed RAG Question-Answering System

A retrieval pipeline over a large biomedical document corpus, with a systematic benchmark of embedding models, vector databases and retrieval strategies. Hybrid search produced the largest gain — not a larger model.

RAGHybrid searchBenchmarkingVector DBs
View on GitHub
M.Sc. Thesis · Ongoing

Simulation-Based Data Generation & ML Model Development

M.Sc. thesis: a physics-based simulation that generates training data, then an ML model trained and iteratively optimised on it — covering the full loop from data generation to evaluation.

SimulationMLPyTorch
(04)Beyond work

The rest of me

Work is the core of this site — but not all of it. A few things I'm building out.

Let's build something.

Whether it's a role, a collaboration or a question about my work — I'd love to hear from you.