high performance computing on graphics processing units: hgpu.org

hgpu.org » Programming » Algorithms » Duplicate Detection on GPUs

Duplicate Detection on GPUs

Benedikt Forchhammer, Thorsten Papenbrock, Thomas Stening, Sven Viehmeier, Uwe Draisbach, Felix Naumann

Hasso Plattner Institute, 14482 Potsdam, Germany

15th GI-Symposium Database Systems for Business, Technology and Web (BTW), 2013

@article{forchhammer2013duplicate,

title={Duplicate Detection on GPUs},

author={Forchhammer, Benedikt and Papenbrock, Thorsten and Stening, Thomas and Viehmeier, Sven and Draisbach, Uwe and Naumann, Felix},

year={2013}

}

Download (PDF)

View

Source

4558

views

With the ever increasing volume of data and the ability to integrate different data sources, data quality problems abound. Duplicate detection, as an integral part of data cleansing, is essential in modern information systems. We present a complete duplicate detection workflow that utilizes the capabilities of modern graphics processing units (GPUs) to increase the efficiency of finding duplicates in very large datasets. Our solution covers several well-known algorithms for pair selection, attribute-wise similarity comparison, record-wise similarity aggregation, and clustering. We redesigned these algorithms to run memory-efficiently and in parallel on the GPU. Our experiments demonstrate that the GPU-based workflow is able to outperform a CPU-based implementation on large, real-world datasets. For instance, the GPU-based algorithm deduplicates a dataset with 1.8m entities 10 times faster than a common CPU-based algorithm using comparably priced hardware.

Tags: Algorithms, Computer science, nVidia, nVidia GeForce GTX 570, OpenCL

March 21, 2013 by hgpu

No votes yet.

Please wait...

Your response

You must be logged in to post a comment.

* * *

high performance computing on graphics processing units: hgpu.org

Duplicate Detection on GPUs

Your response

Recent source codes

UniCoder: Unified Visual-to-Code Generation via Symbolic Rewards and Reference-Guided Code Optimization

CuFuzz: An API-Knowledge-Graph Coverage-Driven Fuzzing Framework for CUDA Libraries

AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning

Probe-and-Refine Tuning of Repository Guidance for AI Coding Agents

CUDAnalyst (CUDA + Analyst)

CodegenBench

KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels

CUDA Kernel Fusion Benchmarks

IntelliKit: Agent-first tooling for AMD hardware

DITRON: Distributed Compiler based on Triton for Parallel Systems

Most viewed papers (last 30 days)

Duplicate Detection on GPUs

Share this:

Your response

Recent source codes

Most viewed papers (last 30 days)