HKR

Name: HKR
Published: 2020-07-07
License: CC-BY-NC-ND-4.0

Handwritten Kazakh and Russian (HKR) Database for Text Recognition

Dataset Information

Modalities

Images, Texts

Languages

Russian, Kazakh

Introduced

2020

License

CC-BY-NC-ND-4.0

Homepage

Official Website

Contents

Overview
Associated Benchmarks
Recent Benchmark Submissions
Research Papers

Overview

The database is written in Cyrillic and shares the same 33 characters. Besides these characters, the Kazakh alphabet also contains 9 additional specific characters. This dataset is a collection of forms. The sources of all the forms in the datasets were generated by LATEX which subsequently was filled out by persons with their handwriting. The database consists of more than 1400 filled forms. There are approximately 63000 sentences, more than 715699 symbols produced by approximately 200 diferent writers. We utilized three different datasets described as following:

Handwritten samples (Forms) of keywords in Kazakh and Russian (Areas, Cities , Village , etc.)
Handwritten Kazakh and Russian alphabet in cyrillic
Handwritten samples (Forms) of poems in Russian

Image source: https://github.com/abdoelsayed2016/HKR_Dataset

Variants: HKR

Associated Benchmarks

This dataset is used in 1 benchmark:

Handwritten Text Recognition - Metrics: CER

Recent Benchmark Submissions

Task	Model	Paper	Date
Handwritten Text Recognition	StackMix+Blots	StackMix and Blot Augmentations for …	2021-08-26

Research Papers

Recent papers with results on this dataset:

StackMix and Blot Augmentations for Handwritten Text Recognition (2021) -

External Links:

HKR

Overview edit

Associated Benchmarks

Recent Benchmark Submissions

Research Papers

Edit Dataset Information

Overview