Algorithmic Reviews

LLM & Probabilistic Approaches

LLMs, Bayesian optimization, uncertainty-aware ML, probabilistic search, and related ML-based decision tools.

All research notes
LLM & Probabilistic Approaches

Denoising Score Matching: Why a Conditional Target Learns the Marginal Score

Diffusion models need the marginal score to reverse a noising process. This note explains why a tractable conditional Gaussian target learns that score and in what sense the two losses are equivalent.

Stanford CME296 Diffusion & Large Vision Models | Spring 2026 | Lecture 2 - Score matching

LLM & Probabilistic Approaches

GraphRAG for Engineering Diagrams: ChatP&ID and P&ID Retrieval

A critical note on ChatP&ID: P&IDs are better treated as structured engineering knowledge graphs than as raw images or raw XML, but the benchmark mainly validates context engineering rather than a new GraphRAG algorithm.

GraphRAG for Engineering Diagrams: ChatP&ID Enables LLM Interaction with P&IDs

LLM & Probabilistic Approaches

Tolerance Ball Acquisition for Specification-Driven Inverse Design

A critical note on Tolerance Ball acquisition: a clean probability-of-feasibility objective for specification-driven inverse design, but not a direct optimizer of diversity, boundary coverage, or global feasible-set recovery.

Range-aware Bayesian optimization for discovering diverse designs within target property windows

LLM & Probabilistic Approaches

BOHB and MILP for Multi-Timescale LH2 Supply Chain Design

A critical note on a BOHB-MILP framework for international liquid hydrogen supply-chain design under hourly renewable variability, weekly shipping, lead time, and sampled demand-weather scenarios.

Techno-economic analysis for design and management of international green hydrogen supply chain under uncertainty: An integrated temporal planning approach

LLM & Probabilistic Approaches

Do Transformers Need Three Projections?

A critical note on Q/K/V projection sharing in Transformer attention, where K=V preserves query-key directionality while cutting KV-cache memory in half.

Do Transformers Need Three Projections? Systematic Study of QKV Variants

LLM & Probabilistic Approaches

LoRA and QLoRA: Low-Rank Adaptation Under Memory Constraints

LoRA reduces fine-tuning cost by learning low-rank task-specific corrections on top of frozen pretrained weights, while QLoRA adds 4-bit quantization so larger base models can be adapted under tighter GPU memory constraints.

Probabilistic Heuristics & Bayesian Search

FlashAttention: Exact Attention as an IO-Aware Streaming Computation

FlashAttention is not an approximation to attention. Its core idea is to avoid materializing the N by N attention matrix in HBM by computing tiled attention in SRAM and maintaining online softmax statistics.

Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 4 - LLM Training