Scalable LLM serving and inference optimization
Currently working on researching and building methods to explore scalable, lower-cost LLM serving and inference optimization.
Independent experiments, collaborative research, and software implementations. Technical briefs document objectives, methods, results, and limitations.
Currently working on researching and building methods to explore scalable, lower-cost LLM serving and inference optimization.
A controlled CIFAR-100 comparison of a residual CNN and a Vision Transformer trained from scratch under matched parameter and estimated compute budgets, measuring accuracy, calibration, noise response, and runtime to inform model selection.
Three MNIST experiments compare reconstruction and subspace geometry, then test architectural constraints and training choices. Includes the full paper and all four original figures.
A video-editing prototype uses speech and semantic scores to plan cuts, then exports a video and an inspectable timeline.
This project extracts actions with supporting source text and prepares a controlled semantic role labeling experiment.
This prototype links caption claims to supplied observations and evidence you can inspect.
This statistical baseline compares incoming text with a reference corpus and reports changes in language.