If you dig through their methodology description what you find is that they're adjusting their numbers based on papers published by people who work for Waymo. And they also promise that they're counting all the accidents where there was any damage or anybody got hurt, because of course they report all of those, doing otherwise would be a crime! Or at least a rule violation. (From their phrasing, I get the impression that there are a large number of unreported, or at least uncounted here, "the Waymo vehicle bumped into something at very low speed and no damage was done" incidents.) And that they're still doing an unspecified amount of remote driving by humans. Which all boils down to "Self-driving car company says its self-driving cars are safer than human drivers!"
I get that the data cannot be usefully compared one to one with the data for human driver. Even if you had perfect reporting, Waymo simply isn't going to choose the same routes human drivers will; there will be stretches of road Waymo will likely never travel, just because of its business model. And there's probably not a lot of good research being done on these statistics (although I don't know that for a fact), so it makes a degree of sense that Waymo employees would be publishing. But claiming "we're using peer reviewed standards for our data adjustments" without a disclaimer that the writers work for Waymo is not a good look, nor does it encourage confidence on my part.
But one item I was not able to find, is where Waymo would fall in a ranking of human drivers, What happens if you take out the worst 1% or 5% of human drivers? How about if you take out incidents involving impaired drivers? It seems like these would be natural data sets, but I was not able to find them. (Maybe my search-fu and time is just insufficient.)