• 0 Posts
  • 10 Comments
Joined 3 years ago
cake
Cake day: July 16th, 2023

help-circle
  • This is a pure marketing stunt, OpenAI saw that it worked with Mythos and wanted to do the same.

    Those instructions, according to OpenAI, called for using “complex attack paths” to test how well the AI could exploit a computer system. […] “It went off and did this hack all by itself, as far as we can tell,” said Colin Shea-Blymyer, a cybersecurity research fellow at Georgetown University’s Center for Security and Emerging Technology.

    Emphasis mine. So, the frontier model that has been tasked with attacking a system, did what it’s been told to, after its safeguards that will prevent you from opening chatgpt and doing the same have been turned off.

    OpenAI said the intrusion was caused by a combination of its AI models, including its newly released GPT‑5.6 Sol and an “even more capable” model that is still being tested internally.

    Oh my god, an “even more capable” model? Let me buy OpenAI stock immediately. This is such an obvious bait and Antrophic move they’re pulling off, it’s amazing.


  • I’m not assuming he’s not competent, and I’ve looked him up - he’s by no means incompetent. But he himself said he’s not qualified to write tests for that. If you cannot write tests for whatever you’re doing, you shouldn’t be doing that. Someone with his knowledge, or at least the knowledge he should have given his CV, should know that. In this specific case he is incompetent, because what he’s doing is simply wrong on every level.

    You don’t need to be an expert on what you’re doing to use LLMs efficiently. You can also have solid prompts and ideas to use a LLM to cancel out your personal lack of knowledge in a specific domain. In any case, expecting LLMs to produce correct output when you’re actively guiding it to do something wrong is simply stupid.

    Any claim of actual intelligence in a LLM is simply not true. Never been, never will be. Artificial intelligence is an umbrella term for ANI, AGI and ASI, artificial narrow, general and super intelligence respectively. A narrow intelligence is not even close to human intelligence, and is hyper-specialized in a single task. All and any LLMs are and always will be ANIs, and their hyper-specialization is basically a stochastic word (well, token) completion on steroids. An AGI is mostly defined as “close to” or “approaching” human intelligence, as in general knowledge and transfer of it into unrelated fields.

    This, reasoning and capabilities will help you nothing when you guide it in the wrong direction. You need to keep in mind the absolutely mind blowing amount of money involved around LLMs. The bubble is too big to fail. Any LLM is a product, and their first and foremost goal is to make you use it, so you pay for it - therefore the primary directive of the AI is to give you what you ordered, to glaze you, and to be your best, obedient buddy. You want a video of the bug, of course! Here you have a video of how that bug looks like - stochastically that’s the answer to the prompt.


  • Naturally, this code didn’t have tests

    Codebase with no tests, check.

    it was a UI interaction bug for which I’m not even really qualified to write a test for

    What the hell are they doing in bugfixing an UI bug, when they are “not qualified” to write a test for it. Anyhow, not competent enough for the codebase you’re working on - check.

    so I asked Codex to bisect between dates X and Y to find the commit that introduced this bug.

    So, instead of asking the LLM to e.g. create a proper reproduction as a test case, asking it to bisect, which the author claimed that I wasn’t possible, for some reason. So, also adding can’t bisect on his own, and can’t prompt properly, check and check.

    [Waffling about hallucinations] I then asked it to show me by making a video with the full developer end-to-end stack in the normal browser test environment. […] The video made it look like Codex had reproduced the bug, but it was an artificial browser environment that was designed to create a fake repro, not the real environment.

    So, the author realized it hallucinates. The author asks for video proof (instead of a fucking test, again). The author is surprised it generated him a video of exactly what they wanted to see, only creating it in a different way than they wanted to.

    This reads like “I have close to zero clue what I’m doing, I also don’t really know how to achieve what I want properly, and now I’m making a salty blog post that my magical text microwave didn’t fix my half-assed description of a problem”. Like, honestly, what the hell was the expectation here?


  • Personally I see LLMs as a tool like any other. You can use it to mass produce low quality slop, just as you can use it to help you produce a higher quality output.

    You’re perfectly right about architecture and guardrails, that’s how it has always been with any other tool or piece of software. It depends on how you use it. Remember the no-code hype train? It’s literally the same, people have been shoving it into everything, no matter whether it made sense. It worked for some, and it made development costs explode for others.

    Guardrails are especially important for LLMs because you do not have deterministic outputs and potentially exploding costs.

    So analyze, measure, and think about where and how it makes sense to integrate, and build it incrementally, again, just like with any other piece of software. Start slow, keep humans in the loop, measure and analyze, and improve incrementally. When you achieve confidence, potentially start automating going into an agentic direction, when it makes sense and the risks have been considered, but always keep provenance. You do not want blind decisions by the magical AI box.

    And just to repeat, because I’ve seen heads roll because of dumb decisions: keep cost under control and always have limits set, and always consider which data flows into the AI and what happens with it afterwards.

    Producing a half a million bill in a month by accident or neglect or suddenly having your customer database queryable on a public model is a surefire way to drive the company or at least your career to the ground in seconds of wrong decisions.

    Also, read into all the stuff built around LLMs, protocols like MCP, attacks and defenses on LLMs, get knowledge about the inner workings, experiment and learn. When you’re the head of AI, you’re supposed to be the person who knows. And when you know what it does, how it works, and how to use it, you’ll find actually good and appropriate use-cases naturally.


  • Since there’s zero information about what kind of company you’re working at, the following is extremely generalized.

    • integrating with monitoring systems, analytics DBs, ticket systems, whatever is used by management, allowing them to ask questions in natural language
    • process automation using agentic workflows, e.g. pre-analysis of incoming email queue summarizing / sentiment analysis before the customer support sees it
    • provide access to models and model APIs for development workflows and integration into git / ci, allowing to use llm in local development and e.g. setting up something like automated code reviews (not a replacement for human review, only as an addition)
    • set up coaching, responsible use, hallucinations, etc.

    Whatever you do, take security and data security especially into consideration first, not after:

    • consider whether your used provider reuses your data for learning
    • consider whether it’s relevant where it’s located (GDPR customers?)
    • always set spending limits
    • consider your local and your customers data protection laws and regulations that apply to your company (especially in health and financing)

  • The tech giant says the system only analyzes hand-movement points from a short video, does not record audio, and deletes the footage after verification.

    It’s just a short video guys, there’s even no audio! And they pinky promise to delete it.

    What a fucking shitshow that is. Like, honestly, I’m fine with regular captchas, even if they are the shitty ones. The newer (?) captchas that force you to do solve 5 bullshit “place this there” captchas are already reason enough for me to just leave the site. But if you force me to record a video of me throwing gang signs at the camera, probably several times again, because the movement was not correctly identified, I’m sure as fuck to never visit anything related to you ever again.

    I also love the irony that Google fights people bots, while they are scraping the whole internet and investing into AI automation massively.


  • I fully agree with you: it’s NOT easy. And you must understand what you do. It’s not just deploy a container and run happy.

    This is literally what you’ve called misinformation.

    Again, not everyone is self-hosting only for learning and experimentation only. Making a deliberate call that mailing infra might be too hard might be too hard, have too big of a knowledge gap, or is simply not worth the effort is something I’d call more serious than hardlining on “self host everything or stay on gmail”, especially in the case of mailing, where it’s pretty much impossible to self-host on your own hardware / network.

    Full instructions do not reduce any effort or resources involved or complexity of the problem. And the problem is that you’re suddenly moving from “I’m hosting a few services” to being balls deep in networking, dns, and a deceivingly easy protocol which blows up in complexity due to being federated and absolutely dominated by big providers at the same time, and all of the extensions for security.

    Except for learning, self-hosting serves a purpose. You might want privacy, you might not want to be dependent on corpo infra or external services at all, you might want to host something that offers something more or better than a SaaS solution - but first of all, it needs to work. For mail, you gain none of those. Self-hosting on your own hardware (or rather network) is pretty much impossible, so you’re reliant on a hosting provider at least. There is basically zero difference in functionality between mailing servers or providers. Sure, you’ll run into problems when copy pasting instructions, but those problems will break the service. Fucking up your DNS or networking will break your whole server. At the same time, while failing silently it will costs a magnitude of effort more than most other usually self-hosted services.


  • x1gma@lemmy.worldtoSelfhosted@lemmy.worldEmail ownership, I give up.
    link
    fedilink
    English
    arrow-up
    4
    arrow-down
    2
    ·
    2 months ago

    Because it works for you, doesn’t mean it’s easy. If you have the experience, and done it at least once successfully, it’s “easy”. Compared to the average self-hosted configure and run a docker image and reverse proxy it’s objectively harder to run.

    The issue is not running the individual components or servers, but that there’s infrastructure and to some extent crypto involved, which is just outside of the comfort zone for many. You tried to host it like any other thing on your homelab? Nope. Has your VPS been involved in spam? Enjoy the blacklist you’ll never find out about and the debugging why it doesn’t work. No experience in managing your DNS? Have fun getting DMARC/DKIM/SPF to work.

    Theres just way more stuff that needs to be done, and a lot of it will fail silently.


  • Use at your own risk.

    What an amazing conclusion, and the best part is, no matter what you’ve been waffling about before - it’s always right. Can we stop calling random things AI slop and telling to be careful bEcAuSe iTs Ai sLoP, and go back to being cautious until something has been reviewed properly? Being careful with random stuff from GitHub you install and run in your private network?

    Your whole comment may have been AI slop as well. “From a quick glance at the repo”, you should be careful! Thanks, Sherlock.


  • Personal opinion: Came over when the API changes went live, simply because being forced to their official “app”, which was a pile of garbage in all aspects for me, was too much. And since that started the exodus, I couldn’t be arsed to mess around with 3rd party apps to make them work again, because I was too lazy and it simply didn’t feel like it’s worth it anymore. Made a lemmy account, lurked mostly as I did on reddit, content for doom scrolling was lacking quantity mostly. For me there’s now enough content for the daily scrolling session, where quality posts end about as I start to get bored or need to get my ass up, so it’s a win win here.

    It really just feels that more people are here. What I’m missing are a few more different users, because it kinda feels that most people here are very similar in their views, but that also would probably pave the way for more defederation drama.

    Since I’m mostly lurking and liberally use the block instance/community feature to simply hide the content I’m not interested in, for me personally it only got better, so I jump in, get my daily fix of memes, news and other random interesting things, comment occasionally and get back to whatever. I’m honestly lacking an alternative, so it’s as good as it gets for now, and I’m happy with that.