AI/ML researcher and engineer focused on reliable LLM and agent systems.
I build real AI systems, study where their apparent reliability breaks, and design evaluation pipelines that make their failures measurable.
Русский · 中文 · GitHub · Telegram · Email
I am a second-year Applied AI student at Novosibirsk State University and an AI/ML engineer working across research and production-oriented systems.
My current focus is reliable AI:
- factual hallucination detection;
- evaluation of LLM agents;
- evidence grounding and safe abstention;
- shortcut learning and distribution shift;
- reproducible applied ML research.
Across my main projects, I study the same underlying problem:
Why do AI systems often look more reliable than they really are, and how can we evaluate them honestly?
|
Detecting factual hallucinations from internal LLM representations. A white-box approach based on hidden-state probing, contrast directions, uncertainty signals, PCA, and a lightweight linear classifier. Highlights
|
Evidence-grounded verification for web procurement agents. The project checks exact product identity, mandatory specifications, provenance, and whether the agent should confirm or abstain. Highlights
|
Diagnosing shortcut learning and evaluation failure under distribution shift. A postmortem of an ad-classification system whose initial metrics did not survive cleaner validation. Highlights
|
| Project | Failure mode | Research question |
|---|---|---|
| Guardian of Truth | A fluent answer may still be factually wrong | Can internal representations expose hallucinations? |
| ProcureTrace | An agent may confidently confirm an unsupported product | Can every decision remain attached to page evidence? |
| SplitShift | A strong validation score may not generalize | Which shortcuts and distribution shifts inflate performance? |
Together, they form one research direction: measuring and improving the reliability of AI systems beyond demo quality and headline metrics.
- LocalScript AI Agent — offline Qwen-based agent for generating and validating Lua scripts in an isolated environment.
- Orange Pi 6 Plus NPU — neural-network inference on ARM64 Linux using a hardware NPU.
- Demand Forecasting Service — PyTorch LSTM, FastAPI backend, React dashboard, and Docker deployment.
- Dion Background Lab — browser-based real-time person segmentation and background replacement.
- RAG and agent systems — FastAPI, LangGraph, ChromaDB, local and hosted LLMs, tool use, and memory.
- 2nd place — Purple Hack 2026, Avito track
- 3rd place — T1 Hackathon
- 3rd place — Cloud.ru Hackathon
- Prize-winning work in factual hallucination detection and AI-agent systems
Research: LLM evaluation, representation analysis, hallucination detection, agent reliability, distribution shift, ablations, error analysis
ML: Python, PyTorch, scikit-learn, CatBoost, pandas, PCA, TF-IDF, feature engineering
LLM systems: LangGraph, LangChain, RAG, embeddings, reranking, MCP, local LLM inference
Engineering: FastAPI, Django, Docker, Docker Compose, pytest, Git, Linux, React
- research internships and laboratories;
- reliable LLM and agent systems;
- applied AI research with real-world failure modes;
- collaborations where experimental rigor and engineering both matter.
- Telegram: @stavrmoris
- Email: s.mariskin@g.nsu.ru
- GitHub: github.com/stavrmoris
