Minimal Perfect Hashing for Go

This library provides Minimal Perfect Hashing (MPH) using the Compress, Hash and Displace (CHD) algorithm.

What is this useful for?

Primarily, extremely efficient access to potentially very large static datasets, such as geographical data, NLP data sets, etc.

On my 2012 vintage MacBook Air, a benchmark against a wikipedia index with 300K keys against a 2GB TSV dump takes about ~200ns per lookup.

How would it be used?

Typically, the table would be used as a fast index into a (much) larger data set, with values in the table being file offsets or similar.

The tables can be serialized. Numeric values are written in little endian form.

Example code

Building and serializing an MPH hash table (error checking omitted for clarity):

b := mph.Builder()
for k, v := range data {
    b.Add(k, v)
}
h, _ := b.Build()
w, _ := os.Create("data.idx")
b, _ := h.Write(w)

Deserializing the hash table and performing lookups:

r, _ := os.Open("data.idx")
h, _ := h.Read(r)

v := h.Get([]byte("some key"))
if v == nil {
    // Key not found
}

MMAP is also indirectly supported, by deserializing from a byte slice and slicing the keys and values.

The API documentation has more details.

Name		Name	Last commit message	Last commit date
Latest commit History 22 Commits
python		python
COPYING		COPYING
README.md		README.md
chd.go		chd.go
chd_builder.go		chd_builder.go
chd_test.go		chd_test.go
slicereader_fast.go		slicereader_fast.go
slicereader_safe.go		slicereader_safe.go

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

python

python

COPYING

COPYING

README.md

README.md

chd.go

chd.go

chd_builder.go

chd_builder.go

chd_test.go

chd_test.go

slicereader_fast.go

slicereader_fast.go

slicereader_safe.go

slicereader_safe.go

Repository files navigation

Minimal Perfect Hashing for Go

What is this useful for?

How would it be used?

Example code

About

Releases

Packages

License

andradeandrey/mph

Folders and files

Latest commit

History

Repository files navigation

Minimal Perfect Hashing for Go

What is this useful for?

How would it be used?

Example code

About

Resources

License

Stars

Watchers

Forks