Enabling fast preemption via dual-kernel support on GPUs
Journal
Proceedings of the Asia and South Pacific Design Automation Conference, ASP-DAC
Pages
121-126
Date Issued
2017
Author(s)
Abstract
To consider QoS for resource-limited mobile systems, we introduce a fast preemption mechanism on GPUs. First, we involve a dual-kernel execution model to support fine-grained preemption, and a resource allocation policy to avoid resource fragmentation problem. Second, we propose a preemption victim selection scheme to reduce the throughput overhead while satisfying a required preemption latency. Evaluations show that we can reach very close to the ideal preemption scheme within 2% difference in terms of deadline violations. Furthermore, on average we improve GPU resource utilization by 2.93× over prior technique during preemption.
Type
conference paper
