Consumer-Rights-Wiki-related ideas

by Jan Malakhovski, version 0.2.0, updated , published , created

My wishlist for Consumer Right Wiki and its sister projects.

Note that each section of this document is a self-contained piece and can be read independently.

Changelog

(Click me to see it.)

v0.2.0 - : continued

v0.1.0 - : init

Table of Contents

(Click me to see it.)

Let’s decide on a standard machine-readable way of recording all URLs of Terms of Service, Privacy Policy, Terms of Use, End-User License Agreements, Warranty Terms, and similar documents relevant to each company and/or product in Consumer Rights Wiki (CRW). That is, let’s decide how to record such things in CRW’s Cargo templates.

“Why?”

For context, on GamersNexus’s YouTube channel, in a relatively recent video titled “We Could Get Sued for This” around t=00:20:00 Steve talks about how Smart TV manufacturers are likely to quietly start changing their Terms of Service agreements in response to GamersNexus’ video titled “216,000,000 Spy TVs | The LG Smart TV Problem” and how web data hoarders in the audience could help GamersNexus by archiving all such agreements by all other Smart TV companies for future research.

So I went to do that, but trying to perform said archivals I discovered that, at least for some companies, even finding the agreements applicable to the devices you are currently using is a non-trivial task! In fact, it appears that even organizations that work in this area for many years, like ToS;DR, miss these things. For example, at the moment of writing of this, their page on Samsung, refers to two documents:

Meanwhile, I found dozens more such documents (or hundreds more, if you count all the references to third-party agreements you are supposedly agreeing to when you agree to their terms of service, in all their per-country and/or per-language versions) by crawling through Google, Bing, and DuckDuckGo results, and by manually crawling around Samsung’s own websites (each of which produced links to never seen before documents!):

And I’m not at all sure that’s all of them.

Per-product sets of applicable legal documents can also be useful in some cases. For example, health-related agreements from the above list don’t apply to all Samsung devices, at least at the moment of writing of this. For another example, ToS;DR rates Open Camera as “Grade C” because its privacy policy honestly mentions that the author’s website uses Google Ad Sense. Even though Open Camera itself deserves “Grade A” since it’s a GPLv3+ app that comes without ads or anything else evil.

Technical implementation

I think the simplest solution here is to add fields containing categorized per-locale legal URLs to https://consumerrights.wiki/w/Template:CompanyCargo and then start referencing them in https://consumerrights.wiki/w/Template:ProductCargo, https://consumerrights.wiki/w/Template:ProductLineCargo, etc.

E.g., in https://consumerrights.wiki/index.php?title=Samsung&action=edit:

{{CompanyCargo
|Description=Large manufacturing conglomerate headquartered in Seoul, South Korea.
|Website=https://samsung.com/
|...
|Legal_General_en_US=https://terms.account.samsung.com/contents/legal/usa/eng/general.html https://www.samsung.com/us/account/privacy-policy/
|Legal_Health=https://samsunghealth.com/terms https://samsunghealth.com/privacy
|Legal_Health_en_US=https://www.samsung.com/us/privacy-policy/consumer-health-data-privacy-statement/
|Legal_TV_en_US=https://www.samsung.com/us/support/legal/LGL10000312/
|...
}}

and then, in https://consumerrights.wiki/index.php?title=Samsung_TVs&action=edit:

{{ProductLineCargo
|Company=Samsung
|LegalCategories=General,TV
|...
}}

and in https://consumerrights.wiki/index.php?title=Samsung_Smartphones&action=edit (does not currently exists):

{{ProductLineCargo
|Company=Samsung
|LegalCategories=General,Health
|...
}}

FULU’s Own Web Archive (on CRW’s Zulip and on CRW)

Let’s create an independent web archive, preferably run by FULU, with the following properties:

“Why?”

For context, on Louis Rossmann YouTube channel, in video from 2024 titled “DCS sues Small YouTuber for accurate product review showing battery issues & misleading warranty” Louis discusses how Deep Cycle Systems (DCS) company changed their battery Warranty Terms on their web site without updating the “last updated” date there, removed the old versions of those pages from the Wayback Machine, and then sued Stefan Fischer, a small YouTuber who reviewed their product, for “defamation” and “misrepresenting their warranty terms”.

As far as I’m aware, that story ends happily for the reviewer, but only because Louis found unedited copies of the offending pages in NLA’s web archive and could demonstrate that DCS were lying and backdating their edited warranty agreements.

Meanwhile, on GamersNexus’s YouTube channel, in a relatively recent video titled “We Could Get Sued for This” around t=00:20:00 Steve talks about how Smart TV manufacturers are likely to quietly start changing their Terms of Service agreements in response to GamersNexus’ video titled “216,000,000 Spy TVs | The LG Smart TV Problem” and how web data hoarders in the audience could help GamersNexus by archiving all such agreements by all other Smart TV companies for future research.

Then, back on Louis Rossmann YouTube channel in his recent video titled “FetchTV says I’m ‘inaccurate & unfair’… you want to play? LET’S GO! 😤” Louis talks about FetchTV set-top box and subscription TV service company that had recently decided to turn older devices made by them defunct and delete all the data stored on them unless their users are already paying them a subscription fee or are willing to pay them a new levy for it, which is not at all what they have promised in their marketing materials when selling those devices. In that video Louis also asks his viewers to go and archive all the relevant documents referenced in his CRW article on the issue personally because they have a tendency of vanishing from the Wayback Machine, as his previous experience shows.

So, wouldn’t it be great if there was a web archiving service that:

“But why aren’t the Wayback Machine and other general-purpose web archives enough?”

In other words, at the moment of writing of this, if you want to legally preserve some evidence against a company you feel might try to screw you over in the future, you can either - archive those pages yourself, - ask Save Page Now and/or a similar feature of another web archiving services to try and archive those pages, hope it works (which it will not for at least some web pages that lazy-load their contents on user inputs) and then hope that the company in question won’t just request removal of those archived snapshots later.

Imagine for a second what would have happened in the DCS case if DCS were more careful and asked to remove their pages not just from the Wayback Machine, but from all public web archives too, which they easily could have done as Wikipedia has a helpful list of all of them:

Needless to say, the latter option is both costly and inapplicable for investigative journalism purposes.

The same observations apply to the FetchTV case. Personally, without all those archived pages I would’ve been gaslighted by their e-mail to Louis. Their arguments there make sense until you look at their original marketing materials.

In other words, the current state of these things is sub-optimal, but a web archive operating as described above would resolve all these issues.

From the legal side of such an operation:

However, while the above ideas appear to be legal to me and, with enough automation, all of the above can be implemented and run by a single person, the unfortunate reality is that you can’t be seen doing such things as a lone computer enthusiast, especially while putting your real name on it. E.g., most notably, see the case of Aaron Swartz. The fact that OpenAI, Anthropic, and other “AI” companies are essentially doing what Aaron did, but not just for published research, which should be in public domain anyway, but for all the data they can get their grabby hands on, on a scale Aaron didn’t even dream of, and now this now appears to be perfectly legal and the federal government wants to invest in them instead of prosecuting them… makes me think really unkind thoughts.

Thus, assuming FULU is a nonprofit (the website does not appear to say either way), such a web archive should probably be run by FULU for legal ass-covering reasons.

Though, of course, all these understandings need to be verified by an actual lawyer.

Technical considerations

From the technical side, if we are to follow the mainstream opinions, then the best current tools for web archival are

However, personally, I think that the WARC format is both overly complicated and also not powerful enough for its stated web archiving purposes:

See WARC’s spec for more info.

The only supposedly good thing about it is that it’s supposedly standard, but it’s a bit of a lie because WARC outputs produced by some tools following the spec fail to be parsed properly by other WARC tools also supposedly following the spec. E.g., notably, data generated by ArchiveBox fails to be parsed by most other WARC tools. See there for discussion.

Additionally, with so many companies getting publicly caught enshittifying their products via their archived web pages I expect many such companies will soon start

thus making their sousveillance using conventional tools as hard as physically possible. After all, many news websites already do this in response to LLM companies scraping them via web archives, see there and there.

Thus, for all of the above reasons, personally, I think other technical solutions are preferable.

Personally, between 2016 and 2022 I’ve been running a personal web archival workflow by making most of my web traffic go through mitmproxy HTTP proxy, with some further processing of the resulting mitmdump files done by custom scripts. However, in 2022 CloudFlare learned to fingerprint and block mitmproxy usage, I assume as a side-effect of them learning to recognize situations where the network-level behaviour of a client does not match its User-Agent, to fight bots. This feature was then copied by most other CDNs, which quickly rendered my mitmproxy setup defunct. So, between 2022 and late 2023 I’ve been trying out various alternative methods and tools for this, eventually becoming dissatisfied with all of them, which is how I ended up developing hoardy-web (also on GitHub, licensed under GPLv3+), which is the tool I use for all my web archival needs now.

So, if it was me implementing this public web archive idea, I would rather simply add the missing bits to hoardy-web because it’s technical design work quite well for a low-maintenance web archiving service because it can generate static website mirrors via its hoardy-web mirror sub-command. hoardy-web mirror produces outputs similar to wget -mpk, but generated offline from already archived data, allowing you to tweak various rendering options without re-downloading anything.

In other words, a web archive that periodically archives a list of URLs and then renders them into dated static website mirrors can be built on top of what hoardy-web already has, plus a tiny harness connecting a browser running hoardy-web’s browser extension to cron or something. Moreover, a service built on top of hoardy-web would only really need to keep its normal GNU/Linux systems and its HTTP daemons up to date. Unlike with all other solutions, hoardy-web itself doesn’t even need to be installed on the machines that would be facing the public, thus reducing possible attack surfaces.