TL;DR
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
OpenDLSS is a public Vulkan reimplementation of NVIDIA’s DLSS 5 Neural Rendering network, with its developer reporting byte-for-byte matches at all 75 tested block boundaries. The project requires users to supply the model weights and compatible NVIDIA hardware; its claims and performance figures come from the project’s GitHub documentation, not an independent evaluation.
A developer has published OpenDLSS, a Vulkan reimplementation of NVIDIA’s DLSS 5 Neural Rendering network, saying it reproduces the reference network’s intermediate results byte for byte across 75 block boundaries. The project makes an implementation available for compatible NVIDIA GPUs, but users must provide the model weights; the supplied report does not establish that NVIDIA has endorsed or independently verified the work.
The project describes the network as a 71-block shifted-window transformer with a global vision transformer at its deepest level, arranged across six pooling levels. It says the network uses E4M3 FP8 activations with FP16 accumulation and has 141 MiB of weights. OpenDLSS takes a rendered frame along with noise, temporal-history information and conditioning values, then produces an RGB residual and a temporal-blend value for each pixel.
According to the repository, the Vulkan implementation runs on NVIDIA Ada or newer GPUs and depends on drivers exposing several specified Vulkan and NVIDIA extensions. Its performance table reports minimum per-frame times on an RTX 4070 SUPER: 2.8 milliseconds at 768-by-768, 7.8 ms at 1920-by-1080, 12.6 ms at 2560-by-1440 and 29.3 ms at 3840-by-2160. The developer says each measurement covers a minimum over 40 frames and notes that GPU clock changes can make medians a few percent higher.
The repository also includes a separate browser-based WebGPU implementation, which the developer says matches the same captures without tensor cores or FP8. Its stated runtime is 72 ms at 512-by-512, compared with 2.7 ms for the Vulkan route at that resolution. The project says its demo places the network in a Filament rendering pipeline, while the command-line tool processes single frames without temporal history. It explicitly says the project does not implement DLSS-SR, a different network.
A Reimplementation Outside NVIDIA’s Stack
OpenDLSS could give graphics developers and researchers a way to inspect and run a claimed reproduction of the neural rendering stage outside NVIDIA’s own software integration. Its published graph, kernels, CPU arithmetic reference and parity tools offer material for testing how this particular network behaves on supported hardware. That may help technical users study implementation choices such as attention, FP8 execution and temporal feedback.
The project is not a general replacement for NVIDIA’s graphics features. The repository says DLSS-SR is not implemented, and the model weights are not supplied with the described setup. Its system requirements also limit use to recent NVIDIA GPUs and particular driver capabilities. The reported timings are project benchmarks, not evidence of performance across other cards, games or production systems.
NVIDIA RTX 4070 SUPER graphics card
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Network Recreates
The project identifies its target as NVIDIA’s DLSS 5: Generative Neural Rendering network, specifically build 310.8.0. It describes the system as re-rendering an image already drawn by a game engine: the network takes a low-dynamic-range proxy frame and additional inputs, then generates image detail and adjusts aspects such as tone, structure and skin according to a style setting. The repository says input and output remain at the same resolution, so this implementation is not an upscaler.
OpenDLSS divides its work between a reference Vulkan route and faster NVIDIA-specific PTX kernels, according to the repository. It also provides a browser WebGPU port as an independent implementation. The project says exactness is checked against fixtures and applies to intermediate block outputs as well as final images. Its demo uses a patched Filament renderer for per-object motion vectors and Vulkan interoperability, while the standalone tool uses single-frame reference captures.
““The intermediates match too, not just the final image: all 75 block boundaries, byte for byte.””
— OpenDLSS project description on GitHub
As an affiliate, we earn on qualifying purchases.
Independent Checks and Model Access
The supplied source is the project’s own GitHub report. It does not identify an independent verification of the byte-for-byte parity claim, benchmark methodology or visual results. The stated performance figures should therefore be read as developer-reported measurements on one named GPU, not as independently reproduced results.
The source says users must supply the weights in a directory with a specified layout, but it does not establish whether those weights are publicly obtainable, what terms govern their use, or whether NVIDIA authorizes their distribution or use. It also does not give a publication date or describe any response from NVIDIA. Those details, along with performance on other hardware and behavior in commercial games, remain unclear.
As an affiliate, we earn on qualifying purchases.
Testing and Project Adoption
The next practical step for interested developers is to inspect the repository’s build instructions, model-directory format and fixture-based parity tools, then test the implementation on hardware and drivers that meet its requirements. The project documents commands for benchmarking, profiling and comparing outputs, but those checks depend on access to the required model files and compatible reference fixtures.
No release schedule, independent audit or NVIDIA response is included in the supplied report. Until those appear, the implementation’s stated exactness and speed remain project claims. Further reporting would need to establish model-weight availability and licensing, reproduce the results on additional systems, and clarify whether the work can be used beyond the project’s demo and test setup.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is OpenDLSS?
It is a GitHub project that describes a Vulkan reimplementation of NVIDIA’s DLSS 5 Neural Rendering network, along with a separate browser WebGPU port.
Does OpenDLSS include the model weights?
No. The project says users must supply the weights in the model-directory format described in its documentation. The supplied report does not say where those weights can be obtained or what terms apply.
Does it upscale images?
The project says it does not. It describes the network as processing a rendered frame and returning an image at the same resolution; it also says DLSS-SR, a different network, is not implemented.
What hardware does the Vulkan version require?
The repository lists Windows and an NVIDIA Ada-generation or newer GPU, plus a driver exposing specified Vulkan and NVIDIA extensions. It reports its performance figures on an RTX 4070 SUPER only.
Has the byte-for-byte match been independently confirmed?
Not in the supplied material. The claim that all 75 block boundaries match byte for byte comes from the OpenDLSS project’s own GitHub description; no independent test is cited.
Source: hn
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
