Open Issues Need Help
View All on GitHub enhancement good first issue Interface/API improvement model training
Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends
Python
#data-matching#data-science#deduplicate-data#deduplication#duckdb#em-algorithm#entity-resolution#fuzzy-matching#record-linkage#spark#uk-gov-data-science
Better error for undialected ColumnExpression 18 days ago
good first issue user experience
Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends
Python
#data-matching#data-science#deduplicate-data#deduplication#duckdb#em-algorithm#entity-resolution#fuzzy-matching#record-linkage#spark#uk-gov-data-science
good first issue
Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends
Python
#data-matching#data-science#deduplicate-data#deduplication#duckdb#em-algorithm#entity-resolution#fuzzy-matching#record-linkage#spark#uk-gov-data-science
Profiling of dates/quantities with a histogram 28 days ago
good first issue profiling
Fast, accurate and scalable probabilistic data linkage with support for multiple SQL backends
Python
#data-matching#data-science#deduplicate-data#deduplication#duckdb#em-algorithm#entity-resolution#fuzzy-matching#record-linkage#spark#uk-gov-data-science