The handwritten trie: indexing electronic ink

Authors:
Walid Aref;Daniel Barbará;Padmavathi Vallabhaneni
Affiliations:
Matsushita Information Technology Laboratory, 2 Research Way, 3rd Floor, Princeton, N.J.;Matsushita Information Technology Laboratory, 2 Research Way, 3rd Floor, Princeton, N.J.;Matsushita Information Technology Laboratory, 2 Research Way, 3rd Floor, Princeton, N.J.
Venue:
SIGMOD '95 Proceedings of the 1995 ACM SIGMOD international conference on Management of data
Year:
1995

Citing 6
Cited 4

The automatic recognition of gestures

The automatic recognition of gestures
Fundamentals of speech recognition

Fundamentals of speech recognition
Pictographic naming

CHI '93 INTERACT '93 and CHI '93 Conference Companion on Human Factors in Computing Systems
The art of computer programming, volume 3: (2nd ed.) sorting and searching

The art of computer programming, volume 3: (2nd ed.) sorting and searching
Trie memory

Communications of the ACM
The Power of PenPoint

The Power of PenPoint

On handling electronic ink

ACM Computing Surveys (CSUR)
SP-GiST: An Extensible Database Index for Supporting Space Partitioning Trees

Journal of Intelligent Information Systems
Ink Retrieval from Handwritten Documents

IDEAL '00 Proceedings of the Second International Conference on Intelligent Data Engineering and Automated Learning, Data Mining, Financial Engineering, and Intelligent Agents
Machine learning in a multimedia document retrieval framework

IBM Systems Journal

Quantified Score

Hi-index	0.00

Visualization

Abstract

The emergence of the pen as the main interface device for personal digital assistants and pen-computers has made handwritten text, and more generally ink, a first-class object. As for any other type of data, the need of retrieval is a prevailing one. Retrieval of handwritten text is more difficult than that of conventional data since it is necessary to identify a handwritten word given slightly different variations in its shape. The current way of addressing this is by using handwriting recognition, which is prone to errors and limits the expressiveness of ink. Alternatively, one can retrieve from the database handwritten words that are similar to a query handwritten word using techniques borrowed from pattern and speech recognition. In particular, Hidden Markov Models (HMM) can be used as representatives of the handwritten words in the database. However, using HMM techniques to match the input against every item in the database (sequential searching) is unacceptably slow and does not scale up for large ink databases. In this paper, an indexing technique based on HMMs is proposed. The new index is a variation of the trie data structure that uses HMMs and a new search algorithm to provide approximate matching. Each node in the tree contains handwritten letters, where each letter is represented by an HMM. Branching in the trie is based on the ranking of matches given by the HMMs. The new search algorithm is parametrized so that it provides means for controlling the matching quality of the search process via a time-based budget. The index dramatically improves the search time in a database of handwritten words. Due to the variety of platforms for which this work is aimed, ranging from personal digital assistants to desktop computers, we implemented both main-memory and disk-based systems. The implementations are reported in this paper, along with performance results that show the practicality of the technique under a variety of conditions.