Computational Linguist · NLP & AI Engineer United States · Open to opportunities

Hasan Can Biyik

Native Turkish speaker with an M.S. in Computational Linguistics. I turn language research into tested, usable systems—from multilingual transformers and LLM evaluation to retrieval agents and deployed ML services.

View selected work → Contact me

Research rigor, practical systems.

I study how meaning changes across languages and contexts, with work on euphemism, humor, sarcasm, irony, and prosody. My peer-reviewed research focuses on Turkish euphemism detection and Turkish–English transfer with multilingual transformers.

I also build end-to-end NLP and AI systems—from multilingual datasets and model evaluation to retrieval services, APIs, testing, and deployment. I’m now bringing that combination of linguistic analysis and engineering to industry roles.

core
python pytorch hugging face nlp / nlu transformers multilingual nlp llm evaluation
ml & data
scikit-learn pandas sql xgboost statistical evaluation data annotation
ai systems
fastapi langgraph rag chromadb streamlit
mlops & infra
docker mlflow prometheus azure kubernetes terraform / aws git

Publications and presentations.

2026 SIGTURK at EACL 2026
Rabat, Morocco
Publication ↗
2024 SIGTURK at ACL 2024
Bangkok, Thailand
Publication ↗
2024 Student Research
Montclair State University
Presentation ↗
First author on both peer-reviewed papers.

Selected work.

01
Agentic discovery API over 2,000 synthetic biospecimen records. LangGraph routes natural-language requests to approved structured and semantic tools, while a deterministic SQLite layer enforces server-side access rules and returns permitted records with validated citations. Includes a 23-scenario evaluation harness and a multi-stage Docker deployment on Azure Container Apps.
LangGraph · FastAPI · BGE embeddings · Chroma · SQLite · Docker · Azure
02
Full-stack RAG workflow for immigration documents with source-grounded, page-cited Q&A; translation drafts and a USCIS-oriented certification template for qualified human review; case timelines; and RFE issue extraction. Supports client-scoped retrieval, multimodal Gemini processing with targeted local fallbacks, and 149 passing tests. Built as a workflow demonstration, not a substitute for legal advice.
FastAPI · React · ChromaDB · Gemini 2.5 Flash · BGE-M3 · OPUS-MT · Docker
03
Context-sensitive euphemism disambiguation with XLM-RoBERTa, trained on 19,490 labeled examples across seven languages and evaluated at 0.808 macro-F1. Provides single and CSV batch inference, language detection, confidence scores, model weights, and a public FastAPI demo. Its separate 22-language synthetic transfer benchmark is presented as preliminary rather than production-readiness evidence.
XLM-RoBERTa · PyTorch · Hugging Face · FastAPI · Docker · GitHub Actions
04
Humor & Sarcasm Evaluation research
Built controlled evaluations for context-sensitive humor and sarcasm judgments across local and frontier models. The humor study produced 270,675 outputs across 21 configurations; the sarcasm benchmark contains 59,698 LLM judgments over 3,142 participant responses. Developed resumable batch pipelines, agreement and context-effect analyses, evidence-linked coding of model explanations, and 43 automated tests.
Python · LLM evaluation · statistical testing · Qualtrics · Ollama · Slurm HPC
05
Predictive Maintenance MLOps reference system
End-to-end MLOps reference implementation on the 10,000-row UCI AI4I dataset, where XGBoost reached 0.768 minority-class F1 versus 0.677 for Random Forest. Tracks and promotes models with MLflow, serves predictions through FastAPI, and adds monitoring, drift-triggered retraining, Docker, Kubernetes, CI/CD, and 24 passing tests. Terraform defines AWS VPC, EKS, and ECR infrastructure; the AWS deployment has not been applied.
XGBoost · MLflow · FastAPI · Docker · Kubernetes · Prometheus · Grafana · Evidently · Terraform
06
Speaker-disjoint speech pipeline for pitch-accent and intonational-boundary detection using 16 acoustic features at 10 ms resolution. A corrected v0.2 evaluation compares classical baselines with Conv1D, CNN, and BiLSTM models on matched tasks; the frame-level Conv1D reached 0.562 prominence F1 and 0.251 boundary F1 in a one-seed, single-held-out-speaker experiment. Packaged as a Python library with a CLI and 36 tests.
Python · audio signal processing · scikit-learn · Conv1D · BiLSTM · pytest
07
Tutorial-inspired research assistant, substantially extended into a multi-service application. LangGraph runs Google, Bing, and Reddit retrieval in parallel, performs per-source GPT analysis, and synthesizes a source-attributed answer. Added a versioned FastAPI backend, Streamlit UI, persistent ChromaDB re-querying, Prometheus metrics, and Docker Compose deployment.
LangGraph · GPT-4 / 3.5 · FastAPI · Streamlit · ChromaDB · Prometheus · Docker Compose
08
Interactive Turkish product-review application integrating a pretrained BERTurk-based binary sentiment classifier with Helsinki-NLP OPUS-MT translation. Supports real-time single-review and multi-review batch inference, exploration of a 235,000-review dataset, and Plotly sentiment and confidence visualizations.
BERTurk · OPUS-MT · Streamlit · Plotly · pandas
09
Local-first résumé–job-description comparison tool combining TF-IDF term coverage with SentenceTransformer similarity. Reports score components, matched and missing terms, per-sentence diagnostics, and a downloadable JSON report without an external inference API after model download. Suggested bullets are explicitly labeled as heuristic drafts whose placeholder metrics must be replaced and verified.
Streamlit · TF-IDF · SentenceTransformers · cosine similarity · local inference
Metrics reflect each project’s documented evaluation setting. Preliminary and pending deployment work is labeled explicitly.

Let’s work together.

I’m especially interested in teams working with language models, multilingual data, evaluation, retrieval, or applied machine learning. If that sounds relevant, I’d be glad to connect.