Artists Built A Site To Escape AI. Scrapers Are Coming For It Anyway

Date:

Share post:

Cara, an image sharing social network built specifically for artists who do not consent to have their work used to train generative AI models, suffered a third security breach in the past ten days, with copyrighted images and metadata from the site appearing on the open source data network Academic Torrents on August 22. This follows previous scrapes of the dataset posted on Reddit (since taken down by the original poster) and the AI model clearinghouse Hugging Face.

“We have been scraped for a third time. This is now a targeted attack meant to cause artists pain,” wrote Cara founder Jingna Zhang in a thread on X Sunday morning. “Scrapers know we’re a volunteer project with no financial means to respond. So they say it’s legal, believing themselves to be untouchable.”

“I believe artists who want to share their work should get to do so without violating the principle of consent. I won’t be bullied into making Cara members-only because some people think ‘consent isn’t always valid.’ You don’t blame victims after harming them.”

She concludes with a plea for community support to raise funds for a legal defense.

Zhang, a fashion and fine art photographer who was named to the Forbes 30 Under 30 list in 2018 and whose work has appeared in Vogue, Elle, Harper’s Bazaar and Time, founded Cara, a public benefit corporation, in late 2022. The network shares some of the same functionality as popular commercial applications like Meta’s Instagram, but has explicit terms of service and features to protect the rights of artists from having their work incorporated into AI training sets without compensation, control or consent.

Artists Seek Refuge from AI

Since it launched, the service has attracted over 1.5 million users, including many professional and hobbyist artists who see AI as a threat to their livelihoods and, often, as a basic affront to the ideals of art. As working artists, they need to share and promote their artwork to get work from clients and collectors, but they do not want their work used to train systems designed to replace them.

When the first wave of generative image LLMs like Midjourney, Stable Diffusion and Dall-E appeared, they were trained used a huge archive of images, including many under copyright, scraped from the Internet by automated bots. Artists realized that participating in public platforms, even their own personal websites, exposed their work to these systems. Zhang built and launched Cara on a shoestring as a place where artists could share work in an environment where no consent could plausibly be implied. While no site can be perfectly secured in today’s online world, it provides a modicum of legal and technological countermeasures and a clear, human-centric ethical stance.

“In a world where it feel like nobody is actually doing anything, Cara is trying to make a difference,” said Zhang in a phone interview. “It might not be perfect, it might not solve all the problems or change how the internet is designed, but we are trying, and that makes a huge difference.”

Scrapers defend their actions

Artists’ claims to own and control their own work online are disputed by individuals, groups and commercial entities who believe that advancing the progress of AI entitles them to any and all data they can obtain, regardless of consent. This was apparently the motivation of the original scraper, who posted an archive of nearly 12 million images from Cara to the sub-Reddit r/DefendingAIArt under the handle “MandarinDrawnPoppy994.”

“Scraping is necessary to develop good models. It’s like building a highway – some houses must be demolished, but in the end everyone benefits,” the poster wrote in a thread titled “[AMA] I scraped all of Cara.”

Zhang says she and others reached out to the original scraper and eventually prevailed on him to take down the post. She adds, “not only did the first scraper delete the dataset, but he has turned around now to offer help, and we’re now co-creating an open source tool separate from Cara that will help people check if they’ve been scraped in new datasets in the future.”

Unfortunately, that was not the end of the problem. Several days later, the site was scraped again by a different actor. This time the data was posted on Hugging Face, a hub of resources for AI developers rooted in the open source community, by a poster under the name “Ioannis/Captive Dreamer.”

After some people reported the post to Hugging Face Trust and Safety, the team responded that “we have reviewed these [copyright reports] carefully. Because no copies of the artworks are hosted here, and because the URLs [in the dataset] point to the copies the artists published on Cara, there is nothing hosted on Hugging Face that we can disable through our notice and takedown process. This is not a judgement about who owns the works (the artists do); it is about what is stored on our servers. Further copyright reports on the same basis will not change this outcome.”

Hugging Face did not respond to a request for further comment for this story.

Third attack in 10 days

Now on Saturday, August 22, Zhang says the site was scraped for a third time, with the perpetrator taking just 123K images, but also a second data set that includes users’ text posts, bio and information they share on the site. That archive has been posted on Academic Torrents, a site that “was established to meet the demands of science in the age of big data” by providing data for researchers, according to its “About” page.

The series of attacks has shaken the members of the community, despite most users recognizing there was not much that the Cara team could do in the face of escalating capabilities of scrapers and hackers.

“I never had an expectation that scraping couldn’t happen,” wrote a poster under the name Venoregard, representing the sentiments of many commenters on the thread. “I was just happy to see an art community free from genAI images so I didn’t have to wonder as a non-artist if something was human made. I mean if you can screenshot something you can probably feed it into AI. Art isn’t unstealable, it’s bound to happen. But Cara isn’t doing the feeding and nor are they hosting genAI images on their site, which is a huge step up from the largest social media sites right now imo.”

Zhang says the incursions not only represent an attack on the values of Cara artists and the entire notion that consent matters, they also drain the company’s scarce financial resources. She says the scrapes mean thousands of extra dollars in server fees to handle the automated traffic, at a moment when Cara was hoping to break even or come out slightly ahead for the year. Zhang says she is looking into doing a GoFundMe to help defray the unforeseen expenses.

Zhang says she believes the two most recent attacks are motivated by animus against the Cara community simply for taking a stand against generative AI in the arts. There is ample evidence to suggest that point of view is fairly widespread in communities like r/DefendingAIArt, who believe themselves persecuted by human artists for using image generation tools and posting the outputs as their personal creations.

“They go through the trouble of harassing us for using AI and making such a big deal about how they don’t want their art being used to train AI, so to get back at them [original poster, MandarinDrawnPoppy994] took their art to use to train models. An eye for an eye type of situation. Doesn’t make it right of course, but that’s the purpose here,” wrote a redditor going by Dragin410 in the AMA thread.

“Somebody really wants to highlight how they can troll us,” she said. “They’re like, ‘you say you want to protect artists, you say you’re opting out. So we’re going to violate your consent, violate your saying no.’ At this point this is nothing but malicious, targeted harassment meant to inflict harm.”

Source link

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Related articles

U.S.-Canada Trade ‘War’ Involves Antlers, Fishing Rods, Wigs And More

ToplineTrade tensions between the U.S. and Canada significantly escalated this weekend when Canadian Prime Minister Mark Carney said...

Russia Never Realized Full Potential Of Syria’s Tartus Naval Base—And Never Will

A Russian ship is pictured at the Russian naval base in the Syrian Mediterranean port of Tartus on...

Stray Kids Tie One Of The Biggest Rock Bands Of All Time

Stray Kids collect a ninth No. 1 on the Billboard 200 as 'This & That' debuts. The K-pop...

What Time Does ‘Lanterns’ Episode 2 Release On HBO And HBO Max?

One of the most highly-anticipated TV shows of the year lands on HBO and HBO Max this weekend....