Gurindar S. Sohi

Vilas Research Professor
John P. Morgridge Professor
E. David Cronon Professor of Computer Sciences
Gurindar S. Sohi

Research

My research has broadly explored how computer architecture can exploit information about program behavior to achieve performance and efficiency that would be difficult to obtain using conventional architectural mechanisms alone.

Much of my early work concerned instruction-level parallelism and high-performance microprocessors. This included dynamically scheduled processors with precise exceptions, high-bandwidth and non-blocking memory systems, multiscalar processing and thread-level speculation, memory dependence prediction, instruction reuse, and trace caches. Several of these ideas subsequently became standard components of commercial high-performance microprocessors, generations of which have collectively sold tens of billions of units. Processors incorporating these and related techniques are used every day in computing devices and systems relied upon by billions of people worldwide, making the ideas part of the foundation of modern computing systems.

One theme that runs through much of this work is prediction and speculation. A processor frequently does not need to know something with certainty before making progress. Instead, it can predict the likely outcome, proceed speculatively, and provide mechanisms for detecting and recovering from an incorrect prediction. Memory dependence prediction is one example: rather than determining the precise addresses of all preceding memory operations before allowing a load to execute, the processor predicts whether a dependence is likely to exist and proceeds accordingly.

Later work investigated multicore processors, speculative and data-driven parallel execution, runtime management of parallelism, virtual memory and caching, instruction delivery, and address translation.

My current interests include architectural and memory-system support for AI inference. Large language models place substantial demands on memory capacity and processor-memory bandwidth, particularly for model weights and KV caches. I am interested in mechanisms that exploit properties of these data and computations to reduce those demands while preserving the numerical values and computations expected by software. More generally, I remain interested in how prediction, speculation, representation, and hardware/software abstractions can be used to overcome emerging performance and memory-system bottlenecks.

Selected Research Contributions

Out-of-order execution and precise state

Early work developed a model for dynamically scheduled high-performance processors that permitted instructions to execute out of program order while maintaining the precise architectural state required for interrupts and exceptions. Variants of this organization became fundamental to modern high-performance microprocessors.

High-bandwidth memory systems

Work on memory systems for superscalar processors demonstrated the importance of allowing multiple cache misses and memory operations to proceed concurrently, helping establish non-blocking caches as an important component of high-performance processors.

Multiscalar processing and thread-level speculation

The Multiscalar project explored speculative execution at a granularity larger than individual instructions, allowing portions of a sequential program to execute concurrently while preserving sequential program semantics.

Memory dependence prediction

Memory dependence prediction showed that processors need not determine exact load and store addresses before deciding whether memory operations could safely execute out of order. Instead, dependence behavior could itself be predicted.

Instruction reuse and value communication

Work on instruction reuse and related mechanisms investigated how repeated computation and dynamically observed values could be exploited to avoid or simplify future computation.

SimpleScalar

Our research group developed SimpleScalar, a processor simulation infrastructure that became widely used in computer-architecture research and education.

Parallel execution and multicore systems

Subsequent research explored speculative and data-driven parallelization, control of excessive parallelism, cooperative caching, reliability, and mechanisms for making effective use of multicore processors.

Instruction delivery and address translation

More recent work has revisited instruction delivery and virtual-memory translation, including instruction pre-sending and program-counter-based assistance for data-address translation.

Architecture for AI inference

Current work examines how architectural and memory-system mechanisms can reduce the storage and data-movement requirements of large-language-model inference while leaving the numerical representation and computations visible to software unchanged.