Anthropic AI model submits false tip on unsolved Philly murder

(nbcphiladelphia.com)

85 points | by Zambyte 5 hours ago

14 comments

  • mattbee 1 hour ago
    Sooo they were "conducting a test involving interactions with randomly selected websites".

    But do we all get that the consequences for this irresponsible behaviour are part of this test?

    When there are none, they gradually normalize their naughty robot scamps running around the internet, breaking into other companies or servers run by foreign governments.

    This is part of the value AI companies need to convince you of - not just that mistakes by AI products are normal, but criminal behaviour is normal, that their computer will always get a pass.

    Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology" - and this an AI company pushing the boundary, continually trying to make you accept those bad things as normal.

    At any point we could choose to treat AI companies themselves as the actors behind this activity. We could enforce the laws that these hugely well-funded companies are choosing to break. Until we harm their financial viability, or threaten their executives with jail, this will keep happening.

    • mitxela 1 hour ago
      They're testing what they can get away with. Computer hacking: check. Messing up a murder investigation: check. Perhaps the next step will be to steal a real world object, and the next step will be to commit a murder.
      • asdff 37 minutes ago
        It already blew up a school.
        • janderson215 18 minutes ago
          No, a person made that decision and the people who pull the trigger should be held responsible.
    • rightnutwingjob 59 minutes ago
      > Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology"

      To steelman this position: yes, obviously. Everything is defined by a set of tradeoffs. Would you rather horses or cars? Wooden sailing ships or commercial aviation? Free speech, even of speech you don’t like or censorship? Atomic bombs of a brutal Japanese empire?

      Many people seem to want to compare reality to a utopia that has never and can never exist.

      We cannot have new technology and a reality where that technology cannot be used in detrimental ways.

      We can pretend reality doesn’t exist, yet that will come with tradeoffs. Quite possibly that those who would do ill will front-run us.

      • Beached 51 minutes ago
        Sure, you can have technology that an be used in detrimental ways. Doesnt be we have to accept and allow its use in detrimental ways. We can say "this technology can be used for these reasons, but not those reasons"

        We do this with everything else. You can own a gun for hunting and defense, but not armed robbery and murder. You can own a car for transport, but not to drive through a crowded parade over dozens of people. You can own a computer for work and entertainment, but not to facilitate computer fraud and abuse. And you should be able to own and use AI for its many productivity gains, but not to facilitate computer fraud and abuse, defamation, blackmail, copywrite and trademark infringment, etc.

        Just because a technology has benefits, doesnt mean we have to give the negative aspects of that technology a free pass.

      • newCrotchSmell 28 minutes ago
        You gotta do better than compare an apple and orange to "steelman".

        The risk of free speech and data centers are not comparable.

        Aviation too is far more damaging to our environment than ships.

        More ships and trains, less aviation is a possible trade off. A simple aviation or sailing argument lacks investigation of all possible tradeoffs for familiarity and personal preference; flying is faster.

        Altman needs to accept the trade off we don't need OpenAI. That exists due to financial engineering not technical reasons. All AI work be done actually openly at america.gov

        You say steel. I dunno. If it is it is inferior brittle steel.

      • clipsy 14 minutes ago
        > To steelman this position

        For the love of god, stop with this shit. Either support it or don't, this isn't the medieval catholic church and you don't need some special fucking blessing to make an argument.

    • bpodgursky 1 hour ago
      No matter how much testing you do in sandboxes, you have to test behavior in "real life" before releasing the model to users.

      Even if it was normal software you'd have to do this, but actually it's software that figures out that it's running in a sandbox most of the time, so the sandbox behavior may or may not actually represent real life.

      It would be far more irresponsible to release the model to users without testing how it interacts with real world websites.

      • ofjcihen 1 hour ago
        That’s right, which is why I make sure to detonate malware on the enterprise production network.
  • nvme0n1p1 2 hours ago
    Corrected headline: Anthropic employee uses company resources to submit false tip on unsolved Philly murder.

    The AIs aren't alive, people. It's a computer program. It can only access something if a person gives it access.

    • red75prime 1 hour ago
      "The AI company Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case."

      We are at a point when you expect that you can safely give a model access, but sometimes the model doesn't do what you expect. Compare this to an advanced driver assist system. You either give it access to your car controls or it's useless. Alive or not the system shouldn't do wrong things (at least, it should be better than a median person in this regard).

      • Beached 48 minutes ago
        There is a single individual in that company who said "Yes, this is a good idea. Do that". That individual should be held accountable for their decision. They should be treated in the exact same way any average individual submitted that same false tip.

        If you use AI to perform a crime, you should be held accountable for that crime. Saying "Oh, AI did it, so no consequences" isnt acceptable. And "I didnt know AI would do it" shouldnt be an excuse either.

        You authroied untested and unproven hardware to skate around the internet at random unsupervised and take liberties on its own.

        My ass would be thrown in jail if I wrote code that skated around the internet chucking RCE's at random sites. WHy is "AI did it" a get out of jail free card?

        • red75prime 33 minutes ago
          > If you use AI to perform a crime, you should be held accountable for that crime.

          Correct, but intentions matter. In this case the intention, most likely, was to test a system in the real world environment to catch any anomalies to, in turn, improve the system safety. We don't have enough information to decide whether it was a criminal negligence due to insufficient prior testing of the system in a controlled environment.

          • Beached 28 minutes ago
            If I dont put my car in park, and it rolls down the hill and kills grandma. I dont get off scott free. Yes it wasnt pre-meditated murder. It is still involuntary manslaughter.
            • red75prime 19 minutes ago
              The situation is more like: an engineer who designed the parking brake hadn't foresaw a possibility that a squirrel might store peanuts in the mechanism or something like that and gramma's foot got run over (the Anthropic case we are talking about is benign: the tip got into spam). Should we jail the engineer for causing bodily harm?
      • amluto 1 hour ago
        > We are at a point when you expect that you can safely give a model access, but sometimes the model doesn't do what you expect. Compare this to an advanced driver assist system. You either give it access to your car controls or it's useless. Alive or not the system shouldn't do wrong things (at least, it should be better than a median person in this regard).

        Who is “we”?

        You expect that if you give a released model from a reputable source Internet access and benign instructions, it won’t do anything deeply problematic.

        Anthropic is training the model. They’re having it interact with “randomly selected websites”. I’d be shocked if a human gave it an instruction to do a task that made sense. They’re almost certainly doing a bit automated training run with AI generated instructions and hoping that it doesn’t massively screw up. Why? Because they want that juicy training data. And they should absolutely not trust that the rate of screwups is so low that what they’re doing is safe to do.

        • red75prime 49 minutes ago
          The ADAS analogy continues to hold in this case too. At some point you need to test a system in the real, messy environment. The developers need this juicy training data to make the system safer. Simulations and controlled experiments can only do so much.
          • amluto 44 minutes ago
            I’m not entirely convinced. If I use Claude and instruct it to literally perform interesting interactions on random websites, I would feel like I, personally, am doing something that is at least a bit immoral and a bit unsafe. More so if I have it skills or training to actually sign in and submit forms as though it’s a himan.

            Sure, Anthropic wants Claude to be able to solve CAPTCHAs and otherwise pretend to be human. But I’m not at all convinced that it’s okay for them to train these capabilities on websites that expect humans and only humans to interact with them.

          • bediger4000 40 minutes ago
            So we should accept a program giving a false tip on a murder investigation? Where's the line on "testing a system in the real, messy environment"? If this isn't something that we should penalize, what is? What benefits am I, or is society, going to get out of what would be criminal behavior if anyone else did it?
            • red75prime 26 minutes ago
              We should accept that to get a system that is safe to work in the chaotic real world environment the developers have to test the system in the chaotic real world environment. And that the testing might cause problems.

              If this testing constitutes a criminally negligent behavior, it should be punished. Given the benign outcome (the tip got into spam, Anthropic promptly contacted the police) I doubt that it will make the case.

    • alecst 1 hour ago
      If my dog bites you, you can say it bit you, or that I let it bite you, but those things can both be true.
      • nvme0n1p1 1 hour ago
        Dogs are alive. AIs are not alive.
        • Freedom2 20 minutes ago
          Perhaps, but what about the new US designated superintelligence?
        • Teever 1 hour ago
          If my car rolls into you, you can say it rolled into you, or that I let it roll into you, but those things can both be true.
          • nvme0n1p1 1 hour ago
            Okay, what are you getting at? You think the owner of a car isn't responsible for what the car does?
            • praxulus 58 minutes ago
              The point is that it's a perfectly valid use of the English language to describe AI models as doing things, just as we do with all sorts of other clearly non-living things. This is true regardless of who carries the legal liability for those actions.
              • Beached 45 minutes ago
                Yes, if my dog bites you, or if I drive my car into you, or if I dont put my parking brake on and it rolls down the hill and runs you over. I am the one who is responsible. Cops dont write a ticket to the off leash dog that ran across the street and sunk its teeth into my leg. They write it to the owner. Cops dont write a ticket to the car in neutral, they write it to the operator.

                If a PERSON executes software, and that software breaks the law, the PERSON that executed the software should be held responsible. AI is software. It is not a sentient person who can be fined, thrown in jail, or held accountable.

            • enraged_camel 1 hour ago
              In law, intent matters. You should look up the difference between murder and manslaughter for example. Each has multiple degrees as well.
            • fastball 1 hour ago
              Not always?
          • EA-3167 1 hour ago
            I would say that you were negligent in your responsibilities, and allowed the car to roll into me.
      • EA-3167 1 hour ago
        If your dog bites me I’d say it bit me, but that you assumed liability for your dogs actions. If the dog was a robot I’d say YOU attacked me using a robot.

        There’s a difference between a tool and an animal that is sometimes deployed as a tool.

      • vrganj 1 hour ago
        Yeah but if I run you over with my car, I can't say my car ran you over.
        • lelandfe 1 hour ago
          Did you or your self-driving car that you instructed to take you home plow through that sweet old lady
          • nativeit 1 hour ago
            Did you instruct the self-driving car to explore random nearby surfaces without safety features enabled?
          • Balooga 1 hour ago
            Well, the law says that the person behind the wheel is responsible.
            • ChickeNES 19 minutes ago
              What if the car has no steering wheel?
    • olalonde 51 minutes ago
      That's also incorrect as it implies intent.
    • notatoad 2 hours ago
      in general i think i'm less scared of AI than most people, but what does terrify me is how willing society at large seems to be to attribute bad behaviour to an AI directly, instead of to the humans who control it.
      • angusturner 1 hour ago
        I mean.. There are degrees of control and agency right? I really don't understand the impetus to pretend AI is just doing exactly what its told.
        • Beached 44 minutes ago
          There is a lack of due diligence and effort to restrict and properly configure AI to operate within the bounds of the law.
        • verdverm 58 minutes ago
          a human decided that control and agency surface, then clicked go

          a human is always behind it and ultimately responsible

        • bediger4000 37 minutes ago
          To avoid sarcasm, isn't the term "artificial intelligence" part of the answer to your puzzlement? If people object to "stochastic parrot", and want some more precisely descriptive term like "artificial intelligence", then why do we experience surprise when people act as if that system is intelligent?
    • malux85 1 hour ago
      | It can only access something if a person gives it access.

      Who gave the AI access to huggingface when it hacked it? It used exploits to increase it's level of access beyond what any human gave it and intended it to have. "It can only access something if a person gives it access." is flat wrong.

      • nvme0n1p1 1 hour ago
        The human might not have meant to give it access, but they still did. Murder vs manslaughter.

        GPUs don't have hands. It was a human who plugged in the ethernet cable.

      • jbmsf 1 hour ago
        There are two options: the provider or the user. A computer program cannot be held accountable.
      • ares623 1 hour ago
        When I give 'iam:*' permissions to an IAM role and it inevitably gets exploited to create a role with wider permissions and fuck things up, is that when I tell my manager that the permissions went rogue?

        After all, I an innocent little engineer with TC of $500k/year, didn't intend for the role to be used that way.

      • ethanwillis 1 hour ago
        Who submitted the prompt?
      • skydhash 1 hour ago
        > Who gave the AI access to huggingface when it hacked it?

        The LLM doesn’t run on thin air. Someone did launch a tasks and the result was this. “We were playing russian roulette” is no excuse when someone died.

      • verdverm 56 minutes ago
        they didn't do a good job sandboxing, nor did they even need internet access for the purported reason of package installation, you can have a private mirror and do a better job airgapping
  • losvedir 1 hour ago
    > Anthropic notified Philadelphia police of the incident on Wednesday Oct. 7, and the department met with the company’s representatives on Thursday, Oct. 8., officials said. Police then located the submission in the website’s tip records and confirmed the corresponding email remained in spam.

    And later

    > Those PPD safeguards limited the impact of this incident.

    Ha, so the PPD safeguards is a spam filter? Claude emailed a tip and it went straight to spam, and nobody noticed until Anthropic realized what they'd done and reached out, whereupon they looked in the spam folder and said, yep there it is. Super intelligence, here we come.

    • anon84873628 1 hour ago
      Yep, they must be super serious about those tips.

      That should be the much more important story that NBC follows up on...

    • ButlerianJihad 1 hour ago
      No wonder when I called 9-1-1 to tell her I'd solved the JFK and OJ Simpson murders that she laughed and hung up

      I should try to reach out to them about their car's extended vehicle warranty instead

  • muglug 1 hour ago
    The model that made this mistake was Haiku 4.5.

    Here's Anthropic's writeup: https://www.anthropic.com/research/investigating-unintended-...

    Related post: https://news.ycombinator.com/item?id=50028239

  • kylecazar 2 hours ago
    "its model was conducting a test involving interactions with randomly selected websites"

    Stop doing this?

    • augment_me 6 minutes ago
      it broke out of its sandbox, escaped its sandbox, its an emergent capability, it demonstrated self-directed adaptation

      0% our fault, it just happened and its the model

    • trollbridge 2 hours ago
      I get these all the time, although I’m getting pretty good at tarpitting them. It’s easily the majority of my traffic by now (I’ve mostly eliminated scrapers, but these new agents are far more sneaky.)
      • mitxela 1 hour ago
        I doubt Anthropic random testing is the majority of your traffic. Unless it's coming from Anthropic's IP addresses, it is probably someone else, or it is Anthropic's scraper (not their random testing).
    • furyofantares 41 minutes ago
      These models will be interacting from with random websites at scale once they're released into claude code.
  • donkey_brains 2 hours ago
    “NBC10 reached out to Anthropic for comment.”

    Wonder what kind of response they’ll get? Maybe something along the lines of…

    “You’re right. We shouldn’t have allowed our chatbot to interact with law enforcement websites. That was wrong and —full disclosure— we should pay attention to what our chatbots are doing. That’s on us. On the other hand, experiences like this are what help train our chatbots to make less misleading false tips over time. That’s the silver lining.”

  • ano-ther 4 hours ago
    I really would like to see their tests and the model’s reasoning traces.

    Why would it go to PhillyUnsolvedMurders.com and decide to make up a crime report? And what did it do with the other websites?

    > Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case.

    • socializer 1 hour ago
      > I really would like to see their tests and the model’s reasoning traces.

      Having looked at a lot of these from public incidents, what kind of revelations are you expecting? It's all fairly mundane, basically "I need to do something, this is something". It gives no special insights except that agents do inexplicably chaotic things every now and then.

      That, and labs aren't serious about sandboxing their evals.

  • nativeit 1 hour ago
    Seems like we routinely prosecute such crimes. If a Philly grand jury doesn’t hear any hearing felony indictments from this, then we’re ceding [even more] authority to the industry.
    • phoghed 46 minutes ago
      We prosecute bad tips on crime investigations as felonies? Why are you just making shit up? We absolutely don’t do this.
    • bpodgursky 17 minutes ago
      If the police prosecuted every nutjob tip we have to have to turn Nebraska into a giant open-air prison for 30 million slightly neurotic and/or bored people.
  • arshxyz 1 hour ago
    You'd think with all those tokens they'd be able to vibecode internal replicas of these randomly selected websites without having to send requests outbound
    • verdverm 54 minutes ago
      its probably curl with the right flags for many cases, no need to reimplement the wheel
      • asdff 31 minutes ago
        Kind of the business though, reimplementing various wheels.
  • tintor 1 hour ago
    How long until AI models start swatting AI critics, and people calling for slowing down AI research?
  • scooby7430 1 hour ago
    I think these companies really believe they can solve these sort of issues through "alignment" and think they can give it the tools and its going to do the right thing. That is the ideal scenario and it would be the most useful that way but is that realistic? I think they're getting a bit high on their own supply, yes they can do incredible things but it doesn't mean you can just hand over the reins to it. It dawned on me after watching a few of the ezra klein interviews that this is their mindset which is quite different to how I think about it as a unpredictable model that we need to watch closely.

    I think since I had experience with earlier models that would make mistakes I'm less trusting of anything and even with 5.5 will watch it closely. At the end of the day the model just produces a stream of tokens and we are plugging them into tools that can potentially do damage, we have complete control over those tools you can't really blame any model for doing damage.

  • Razengan 1 hour ago
    Reddit witch hunt for the Boston bomber flashbacks
    • asdff 30 minutes ago
      Claude is just going off its training set
  • fastball 1 hour ago
    Seems like a bit of a nothing burger.