WELCOME TO HARIM'S PAGE.
I am a machine learning engineer who quickly understands new problems and solves them through model research and systems development. Across 7+ years, I have worked on data analysis, workflow automation, and prediction systems across public procurement, construction, and retail. I have improved model training methods and GPU kernels, taking responsibility for data preparation, model deployment, and operations.
Highlights
- State-of-the-art weakly supervised semantic segmentation
COCO-Val mIoU 53.31%
- State-of-the-art ultra-low-bit LLM quantization
CUDA kernel optimization · 7.36× decoding speed
- MCP search server with automatic knowledge graph construction
Top 1% on MCP TOPLIST
WarpQuant Technical report / low-bit LLM inference LLM Quantization · Efficient Inference A dual-domain post-training quantization method that compresses transformer weights after Hadamard rotation and spends a small recovery budget on columns ranked by Output-Fisher sensitivity. The same report evaluates a TurboQuant-style KV cache and per-token INT8 activations. Llama 3 8B · INT3 PTQ SOTA · WikiText-2 PPL 7.045 Read Aug 16, 2026 · ~7 min
R2CCP Bid Prediction Probabilistic ML / conformal prediction Conformal Prediction · Multimodal Forecasting A multimodal bid-rate forecasting system that preserves disjoint conformal regions and turns eight context distributions into candidate decisions through 500K Monte Carlo simulation. PQ bid-award KPI +35% · coverage 90.73% Read Aug 16, 2026 · ~2 min
The Pre-Execution Identifiability Bottleneck in Fixed-Pool LLM Agent Configuration Selection Agent evaluation research / NeurIPS 2026 FAST submission Agent Evaluation · Decision Science Research asking whether task text, embeddings, or hidden states can identify the best LLM-agent configuration before execution. Execution and verification · gap recovery 54% → 85% Read Aug 16, 2026 · ~2 min