by Jan Malakhovski, version 0.2.0, updated , published , created
Note that each section of this document is a self-contained piece and can be read independently.
(Click me to see it.)
Let’s decide on a standard machine-readable way of recording all URLs of Terms of Service, Privacy Policy, Terms of Use, End-User License Agreements, Warranty Terms, and similar documents relevant to each company and/or product in Consumer Rights Wiki (CRW). That is, let’s decide how to record such things in CRW’s Cargo templates.
For context, on GamersNexus’s YouTube channel, in a relatively recent video titled “We Could Get Sued for This” around t=00:20:00 Steve talks about how Smart TV manufacturers are likely to quietly start changing their Terms of Service agreements in response to GamersNexus’ video titled “216,000,000 Spy TVs | The LG Smart TV Problem” and how web data hoarders in the audience could help GamersNexus by archiving all such agreements by all other Smart TV companies for future research.
So I went to do that, but trying to perform said archivals I discovered that, at least for some companies, even finding the agreements applicable to the devices you are currently using is a non-trivial task! In fact, it appears that even organizations that work in this area for many years, like ToS;DR, miss these things. For example, at the moment of writing of this, their page on Samsung, refers to two documents:
Meanwhile, I found dozens more such documents (or hundreds more, if you count all the references to third-party agreements you are supposedly agreeing to when you agree to their terms of service, in all their per-country and/or per-language versions) by crawling through Google, Bing, and DuckDuckGo results, and by manually crawling around Samsung’s own websites (each of which produced links to never seen before documents!):
And I’m not at all sure that’s all of them.
Per-product sets of applicable legal documents can also be useful in some cases. For example, health-related agreements from the above list don’t apply to all Samsung devices, at least at the moment of writing of this. For another example, ToS;DR rates Open Camera as “Grade C” because its privacy policy honestly mentions that the author’s website uses Google Ad Sense. Even though Open Camera itself deserves “Grade A” since it’s a GPLv3+ app that comes without ads or anything else evil.
I think the simplest solution here is to add fields containing categorized per-locale legal URLs to https://consumerrights.wiki/w/Template:CompanyCargo and then start referencing them in https://consumerrights.wiki/w/Template:ProductCargo, https://consumerrights.wiki/w/Template:ProductLineCargo, etc.
E.g., in https://consumerrights.wiki/index.php?title=Samsung&action=edit:
{{CompanyCargo
|Description=Large manufacturing conglomerate headquartered in Seoul, South Korea.
|Website=https://samsung.com/
|...
|Legal_General_en_US=https://terms.account.samsung.com/contents/legal/usa/eng/general.html https://www.samsung.com/us/account/privacy-policy/
|Legal_Health=https://samsunghealth.com/terms https://samsunghealth.com/privacy
|Legal_Health_en_US=https://www.samsung.com/us/privacy-policy/consumer-health-data-privacy-statement/
|Legal_TV_en_US=https://www.samsung.com/us/support/legal/LGL10000312/
|...
}}
and then, in https://consumerrights.wiki/index.php?title=Samsung_TVs&action=edit:
{{ProductLineCargo
|Company=Samsung
|LegalCategories=General,TV
|...
}}
and in https://consumerrights.wiki/index.php?title=Samsung_Smartphones&action=edit (does not currently exists):
{{ProductLineCargo
|Company=Samsung
|LegalCategories=General,Health
|...
}}
Let’s create an independent web archive, preferably run by FULU, with the following properties:
it should support easy import and hoarding of snapshots from other web archives and from independent data hoarders, i.e.:
it should be easy to ask it to slurp a list of Wayback Machine URLs into it;
it should be possible for third-parties to submit ZIPs, SingleFile outputs, WARCs (spec), HARs (spec), WRRs, and other similar web page snapshot formats into it;
while it could hoard snapshots of arbitrary web pages, for legal reasons discussed below it would only publicly re-distribute snapshots of consumer-rights-related web pages; which is to say, the list of URLs it would be re-distributing would be vetted;
ideally, it should be used for archival of all the web pages cited by Consumer Rights Wiki (CRW);
at the very least, it should be used for archival of Terms of Service and similar URLs recorded in CRW Cargo templates;
it could also hoard archives of various relevant marketing materials and product listing pages;
but, again, it should not be used a general purpose web archiving service;
it would comply with content removal and DMCA take-down requests from companies legally “owning” those documents by hiding them from public view, as the law requires and like most other web archives do, but, importantly, it would also add any companies making such requests to its watch-list of “companies apparently planning to do something nefarious”, which it would then actively monitor for enshittification, creating CRW incidents and press releases as soon as any enshittification is discovered, and immediately re-publishing all relevant snapshots previously hidden from public eyes, now under the “newsworthy” fair use copyright exception.
For context, on Louis Rossmann YouTube channel, in video from 2024 titled “DCS sues Small YouTuber for accurate product review showing battery issues & misleading warranty” Louis discusses how Deep Cycle Systems (DCS) company changed their battery Warranty Terms on their web site without updating the “last updated” date there, removed the old versions of those pages from the Wayback Machine, and then sued Stefan Fischer, a small YouTuber who reviewed their product, for “defamation” and “misrepresenting their warranty terms”.
As far as I’m aware, that story ends happily for the reviewer, but only because Louis found unedited copies of the offending pages in NLA’s web archive and could demonstrate that DCS were lying and backdating their edited warranty agreements.
Meanwhile, on GamersNexus’s YouTube channel, in a relatively recent video titled “We Could Get Sued for This” around t=00:20:00 Steve talks about how Smart TV manufacturers are likely to quietly start changing their Terms of Service agreements in response to GamersNexus’ video titled “216,000,000 Spy TVs | The LG Smart TV Problem” and how web data hoarders in the audience could help GamersNexus by archiving all such agreements by all other Smart TV companies for future research.
Then, back on Louis Rossmann YouTube channel in his recent video titled “FetchTV says I’m ‘inaccurate & unfair’… you want to play? LET’S GO! 😤” Louis talks about FetchTV set-top box and subscription TV service company that had recently decided to turn older devices made by them defunct and delete all the data stored on them unless their users are already paying them a subscription fee or are willing to pay them a new levy for it, which is not at all what they have promised in their marketing materials when selling those devices. In that video Louis also asks his viewers to go and archive all the relevant documents referenced in his CRW article on the issue personally because they have a tendency of vanishing from the Wayback Machine, as his previous experience shows.
So, wouldn’t it be great if there was a web archiving service that:
would just automagically archive all things referenced by CRW, including by duplicating snapshots from Wayback Machine’s and other similar web archiving services;
where any independent reviewer could just submit URLs of any relevant evidence pages before publishing their reviews, thus effectively making the DCS bullying strategy defunct;
where any interested bystander could help by submitting more relevant URLs and/or their own snapshots.
As far as I understand, the Wayback Machine and similar web archiving services exist in a kind of legal gray zone that did not shrink into nothingness only because
some of what they do falls under “nonprofit and educational” clauses of fair use copyright exceptions and Article 108(h);
outside of that, the Internet as a whole is younger than current copyright term length, so most of the material such archives store and distribute is in violation of copyright law as it currently written, so such archives have to employ the following defensive strategies:
they have to be very cooperative with copyright title holders and domain host owners and hide their archived data from the public on request quite easily (e.g., https://help.archive.org/help/how-do-i-request-to-remove-something-from-archive-org/, https://trove.nla.gov.au/help/categories/websites-category, etc);
they have to respect robots.txt files and declare themselves in User-Agent strings, thus giving domain owners all the chances to block them from archiving things;
they insist that they are “libraries” even though their operations violate “library physics” of “we only have N copies we can lend” all the time;
which is also the place where they get sued the most, like with https://en.wikipedia.org/wiki/Hachette_v._Internet_Archive;
lawyers use such web archiving services and would rather like keep them around too.
The only public web archive known to me that actively resists removal of old snapshots from itself and ignores robots.txt is https://archive.today/, which is why it’s usually classified as a “pirate” web site, FBI subpoenaed its domains before, and their current set of domains is banned in many jurisdictions, see https://en.wikipedia.org/wiki/Archive.today#History.
But, importantly for this discussion, https://archive.today/ sometimes substantially edit archived pages they serve to the public, which is one of the reasons why Wikipedia banned their usage on Wikipedia pages, see https://en.wikipedia.org/wiki/Archive.today and https://en.wikipedia.org/wiki/Wikipedia:Requests_for_comment/Archive.is_RFC_5#Evidence_of_altering_snapshots.
In other words, at the moment of writing of this, if you want to legally preserve some evidence against a company you feel might try to screw you over in the future, you can either - archive those pages yourself, - ask Save Page Now and/or a similar feature of another web archiving services to try and archive those pages, hope it works (which it will not for at least some web pages that lazy-load their contents on user inputs) and then hope that the company in question won’t just request removal of those archived snapshots later.
Imagine for a second what would have happened in the DCS case if DCS were more careful and asked to remove their pages not just from the Wayback Machine, but from all public web archives too, which they easily could have done as Wikipedia has a helpful list of all of them:
Firstly, everyone could have been gaslighted into believing DCS’ claims.
Secondly, even if the original reviewer successfully resisted the gaslighting (e.g., by archiving some of the original materials themselves), to defend themselves against DCS they would have needed some independently collected evidence, which means they would have needed to counter-sue DCS, sue the original domain host owner where those pages were stored as the part of their lawsuit against DCS, and then subpoena all web archives for their copies of those pages, including the ones hidden from the public.
Needless to say, the latter option is both costly and inapplicable for investigative journalism purposes.
The same observations apply to the FetchTV case. Personally, without all those archived pages I would’ve been gaslighted by their e-mail to Louis. Their arguments there make sense until you look at their original marketing materials.
In other words, the current state of these things is sub-optimal, but a web archive operating as described above would resolve all these issues.
From the legal side of such an operation:
From a copyright standpoint, I think it would be quite hard to argue in court that a Terms of Service, a product listing, or a marketing material is “subject to normal commercial exploitation” like a movie or a song is, but it would be quite easy to argue that all such materials are designed for as wide distribution as possible anyway, are “nonprofit and educational”, with some of them also being “news”, as noted above.
While some argue that when Terms of Service of a website forbid its archival, then to ignore those them and archive it anyway is a violation of Computer Fraud and Abuse Act (CFAA), which would make some of the above operations illegal.
However, I don’t think this interpretation makes physical sense: CFAA forbids “access” to computer systems “without authorization”, and the “access” part happens while you are fetching the data, not while you are saving it to disk. So, in theory, if you are not “hacking” anything while accessing the data, you should be in the clear.
E.g., I don’t think that downloading a YouTube video is a violation of CFAA even if YouTube’s Terms of Service forbid it, which, of course, they do, to quote https://www.youtube.com/t/terms?hl=en:
The following restrictions apply to your use of the Service. You are not allowed to:
- access, reproduce, download, distribute, transmit, broadcast, display, sell, license, alter, modify or otherwise use any part of the Service or any Content except:
- as specifically permitted by the Service;
- with prior written permission from YouTube and, if applicable, the respective rights holders; or
- as permitted by applicable law;
However, while the above ideas appear to be legal to me and, with enough automation, all of the above can be implemented and run by a single person, the unfortunate reality is that you can’t be seen doing such things as a lone computer enthusiast, especially while putting your real name on it. E.g., most notably, see the case of Aaron Swartz. The fact that OpenAI, Anthropic, and other “AI” companies are essentially doing what Aaron did, but not just for published research, which should be in public domain anyway, but for all the data they can get their grabby hands on, on a scale Aaron didn’t even dream of, and now this now appears to be perfectly legal and the federal government wants to invest in them instead of prosecuting them… makes me think really unkind thoughts.
Thus, assuming FULU is a nonprofit (the website does not appear to say either way), such a web archive should probably be run by FULU for legal ass-covering reasons.
Though, of course, all these understandings need to be verified by an actual lawyer.
From the technical side, if we are to follow the mainstream opinions, then the best current tools for web archival are
However, personally, I think that the WARC format is both overly complicated and also not powerful enough for its stated web archiving purposes:
See WARC’s spec for more info.
The only supposedly good thing about it is that it’s supposedly standard, but it’s a bit of a lie because WARC outputs produced by some tools following the spec fail to be parsed properly by other WARC tools also supposedly following the spec. E.g., notably, data generated by ArchiveBox fails to be parsed by most other WARC tools. See there for discussion.
Additionally, with so many companies getting publicly caught enshittifying their products via their archived web pages I expect many such companies will soon start
thus making their sousveillance using conventional tools as hard as physically possible. After all, many news websites already do this in response to LLM companies scraping them via web archives, see there and there.
Thus, for all of the above reasons, personally, I think other technical solutions are preferable.
Personally, between 2016 and 2022 I’ve been running a personal web archival workflow by making most of my web traffic go through mitmproxy HTTP proxy, with some further processing of the resulting mitmdump files done by custom scripts. However, in 2022 CloudFlare learned to fingerprint and block mitmproxy usage, I assume as a side-effect of them learning to recognize situations where the network-level behaviour of a client does not match its User-Agent, to fight bots. This feature was then copied by most other CDNs, which quickly rendered my mitmproxy setup defunct. So, between 2022 and late 2023 I’ve been trying out various alternative methods and tools for this, eventually becoming dissatisfied with all of them, which is how I ended up developing hoardy-web (also on GitHub, licensed under GPLv3+), which is the tool I use for all my web archival needs now.
So, if it was me implementing this public web archive idea, I would rather simply add the missing bits to hoardy-web because it’s technical design work quite well for a low-maintenance web archiving service because it can generate static website mirrors via its hoardy-web mirror sub-command. hoardy-web mirror produces outputs similar to wget -mpk, but generated offline from already archived data, allowing you to tweak various rendering options without re-downloading anything.
In other words, a web archive that periodically archives a list of URLs and then renders them into dated static website mirrors can be built on top of what hoardy-web already has, plus a tiny harness connecting a browser running hoardy-web’s browser extension to cron or something. Moreover, a service built on top of hoardy-web would only really need to keep its normal GNU/Linux systems and its HTTP daemons up to date. Unlike with all other solutions, hoardy-web itself doesn’t even need to be installed on the machines that would be facing the public, thus reducing possible attack surfaces.