Technology

34865 readers

40 users here now

This is the official technology community of Lemmy.ml for all news related to creation and use of technology, and to facilitate civil, meaningful discussion around it.

Ask in DM before posting product reviews or ads. All such posts otherwise are subject to removal.

Rules:

1: All Lemmy rules apply

2: Do not post low effort posts

3: NEVER post naziped*gore stuff

4: Always post article URLs or their archived version URLs as sources, NOT screenshots. Help the blind users.

5: personal rants of Big Tech CEOs like Elon Musk are unwelcome (does not include posts about their companies affecting wide range of people)

6: no advertisement posts unless verified as legitimate and non-exploitative/non-consumerist

7: crypto related posts, unless essential, are disallowed

founded 5 years ago

MODERATORS

[email protected]

142

AI is fundamentally 'a surveillance technology' (techcrunch.com)

submitted 1 year ago by [email protected] to c/[email protected]

28 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] [email protected] -3 points 1 year ago* (last edited 1 year ago) (1 children)

on a book is pirating said book.

If the source is literally a piracy website that serves up applications on how to remove DRM from ebooks, it's absolutely piracy. You can't just deny the source and be like "it's not piracy!" The way the data came into your hands was illicitly, not legally. Especially if DRM has been circumvented and removed before it came into your hands.

They didn't go out and buy copies of thousands of books.

Pretty amusing that you think scraping published data somehow constitutes surveillance, though.

I don't, I was making a point about how absurdly large the language models have to be, which is to say, if they have to have that much data on top of thousands of pirated books, it means they fundamentally cannot make the models work without also scraping the internet for data, which is surveillance.

[–] [email protected] 8 points 1 year ago (1 children)

If the source is literally a piracy website that serves up applications on how to remove DRM from ebooks, it’s absolutely piracy. You can’t just deny the source and be like “it’s not piracy!”

They didn’t go out and buy copies of thousands of books.

And if they went to a library and scanned all the books?

I don’t, I was making a point about how absurdly large the language models have to be, which is to say, if they have to have that much data on top of thousands of pirated books, it means they fundamentally cannot make the models work without also scraping the internet for data, which is surveillance.

I mean, it's just not surveillance, by definition. There's no observation, just data ingestion. You're deliberately trying to conflate the words to associate a negative behavior with LLM training to make your argument.

I really don't get why LLMs get everybody all riled up. People have been running Web crawlers since the dawn of the Web.

[–] [email protected] 2 points 1 year ago (1 children)

There's no observation, just data ingestion.

The AI literally observes the training data

[–] [email protected] 0 points 1 year ago

Insofar as my computer observes the data on my hard disk. But I suspect you know what I meant.