0
votes

Any idea about how to get common keys from large set of unsorted_multimap ??? I use file_name(string) as a key and its size(int) as a value. Basically I am scanning a directory for searching duplicate files using boost and holding entry of each file in unsorted_multimap. Once this map is ready I need to output common keys(file_name) and there sizes as a list of duplicate files.

1
Lots of ideas. Sorting them, or using a unique container to track collisions, would be a good start. What did you try? - Useless

1 Answers

0
votes

How to find common keys of an unsorted_multimap ?

The following code searches for a specific filename, and iterates through all elements with the same key:

std::unordered_multimap<std::string, int> mymulti;      // key: filename, value: size  
//... fill the multimap
for (auto x = mymulti.find("fileb"); x != mymulti.end() && x->first == "fileb"; x++) { 
    std::cout << x->second << " ";    // do something !
}
std::cout << "}\n";  // end something !  

How to iterate through an unsorted_multimap, goupring processing by common keys ?

The following code iterates trhough the whole map, and for eacuh key, processes in a subloop the related values:

for (auto i = mymulti.begin(); i != mymulti.end(); ) {  // iterate through multimap
    auto group = i->first;         // start a new group
    std::cout << group << "={";    // start doing something for the group
    do {                    
        std::cout << i->second << " ";  // do something for every values of the group
    } while (++i != mymulti.end() && i->first == group);  // until we change value
    std::cout << "}\n";                 // end something for the group 
}
// end overal processing of the map 

How to find duplicate files (same key and same value ) ?

Using the building blocks above, you could for every filename, you could create a temporary unsorted_map with the size as value, looking if the element is already in the temporary map (duplicate) or adding it (non duplicate).

If the whole purpose of your unsorted_multimapis to process these duplicates, then it would be pbetter, from the start to build an unosorted_map with filenames as keys, and value a multimap with size as sorted key and values, the other elements you collect on the file (full url ? inode ? wathever):

unsorted_map<string, multimap<long, filedata>> myspecialmap;