Back to selected work
CASE STUDY / Deterministic workflows

AI-Powered Technical Interview Assistant

Personalized mock interviews with timed questions, guarded session states, and evidence-based feedback. An eight-state workflow keeps the session predictable while two-stage evaluation separates scoring from coaching.

CONTRIBUTION

Project work · CV and public repository

PROJECT STATUS

Public project · see repository for latest activity

EXPLOREView source on GitHub
PythonChainlitPydanticPlotlyOllama
02 / CONTROL & EVALUATION+
Simplified system architecture
THE SYSTEM, AT A GLANCE

A little less guesswork. A lot more structure.

  • Eight guarded interview states
  • Seven question types
  • Two-stage rubric evaluation
01 / CONTEXT

The problem and its constraints.

A mock interview needs more than generated questions. It needs clear session state, timing, scoring rules, and feedback that a candidate can use. This assistant combines LLM-generated content with explicit Python control flow.

  • Multi-turn sessions need valid state transitions and timing.
  • Scores can drift or inflate without a stable rubric.
  • Provider failures should not make the interview unusable.
02 / CONTRIBUTION

What the evidence supports.

The CV describes Shaswot's work on the eight-state workflow, deterministic scoring and question distribution, multi-provider routing, and typed scorecards. The linked repository documents the interview behavior and evaluation pipeline.

03 / ENGINEERING DECISIONS

Inside the approach.

01

Make the session state explicit

The workflow moves through IDLE, ONBOARDING, GENERATING, INTERVIEWING, EVALUATING, FEEDBACK, COMPLETED, and DEBRIEF. Guarded transitions and server-authoritative timers keep session control separate from generated text.

02

Score first, coach second

The first evaluation stage produces rubric-based scores with evidence; the second generates coaching feedback. Score caps and rubric anchoring are intended to reduce score inflation rather than treating fluent feedback as proof of a good answer.

03

Keep the results inspectable

Pydantic scorecards combine structured feedback with deterministic statistics. Plotly reports and PDF/Markdown exports provide a record that can be reviewed after the interview.

04 / EVALUATION & RELIABILITY

Results with context.

7

Question types

Supported formats described in the project README.

19

Competency dimensions

Scoring dimensions reported in the CV.

The repository documents retries, fallback questions, role validation, and guarded state changes. The CV reports more than 100 tests. Multi-provider routing uses local Ollama validation and OpenAI-compatible providers for evaluation.

05 / ENGINEERING TAKEAWAY

What this project demonstrates.

An LLM can generate and evaluate content while explicit code remains responsible for timing, state, and the shape of the result.

A POSSIBLE NEXT STEP

A useful next step would be a human-rated benchmark of score consistency across providers. This is a proposed evaluation, not an existing result.

Explore the evidence.

Project repository Résumé and reported results
NEXT CASE STUDY

Credit Risk Pipeline