Facial recognition, gait recognition etc. (i.e. a different kind of machine learning) is probably more important for surveillance, but LLMs can play a role, too. Websites, unencrypted messengers, email servers etc. have been collecting craptons of text data, but before LLMs it was fairly difficult to analyse all of it if you don’t have a place to start like “user username with IP address xxx.yyy.zzz.pp called the president a butthole on dd/mm/yyyy”. You could always filter for keywords, but people get creative about it, which is substantially harder when LLMs are used for analyzing. So yeah, “nasty data processing tools”, and that’s actually a big issue because of how nicely it slots into the network of mass data collection that companies have been building for targeted advertising (or so they say, to me it was always clear that a lot of them were already using it for more nefarious purposes).
And e.g. I assume that you can identify most people just by their writing style across tens of otherwise unrelated accounts.
Also, it’s not like you can’t use the hardware that’s training or running LLMs for other kinds of machine learning.
Facial recognition, gait recognition etc. (i.e. a different kind of machine learning) is probably more important for surveillance, but LLMs can play a role, too. Websites, unencrypted messengers, email servers etc. have been collecting craptons of text data, but before LLMs it was fairly difficult to analyse all of it if you don’t have a place to start like “user username with IP address xxx.yyy.zzz.pp called the president a butthole on dd/mm/yyyy”. You could always filter for keywords, but people get creative about it, which is substantially harder when LLMs are used for analyzing. So yeah, “nasty data processing tools”, and that’s actually a big issue because of how nicely it slots into the network of mass data collection that companies have been building for targeted advertising (or so they say, to me it was always clear that a lot of them were already using it for more nefarious purposes).
And e.g. I assume that you can identify most people just by their writing style across tens of otherwise unrelated accounts.
Also, it’s not like you can’t use the hardware that’s training or running LLMs for other kinds of machine learning.