Depends on the harness too. The MS copilot 365 chat interface happens to parse PDF and docx files before handing them to the model. Absolutely fantastic if your PDF is 1300 pages of concatenated reports and you want to find a single detail in it that you can't easily ctrl+f for.
Plenty of code is not publically available, and yet that has not stopped them from RLHF'ing good coding bots. You can surely see the parallels to any other field including finance.
Lots of bad(internal) code isn't accessible, the good open source stuff is quite the contrary is true in finance, lots of amateurs are sharing their thoughts where as actual deep analysis is kept to internal presentations.
My vague understanding of the context window limitation is that it is largely a constraint of the model architecture. So maybe they have special extra long ctx, but it might just be a hard limit of the model itself.
Yes. If they could prove where the de-identified data came from then it wouldn't be de-identified. There's a whole field of statistics dedicated to this problem and often applied to things like national census data.
I am absolutely struggling to sell my workplace (which is entirely knowledge work) on the usefulness of LLMs for proofreading let alone on automation of hairy parts of our workflows. So yeah, even people who should be able to see what is coming are not looking.
Our accountant told me 'he's not letting go of his claude subscription' followed by a long list of things it does for him. And my friend 'nah haven't really used AI' before his description of it clarified he still thinks they are GPT 3.5 chat bots.
The world has changed in 12 months and many people didn't notice, and those that did are eating their peers' lunch.
It's actually a really interesting question to reckon with - can an AI help raise a child and is that good or bad? I suspect there are ways you can use the current frontier that elevate the child's experience, especially when it comes to education and creative play. On the flipside using one of these models to automate storytelling to your 3 year old is probably a bad idea... The tricky bit is where to draw the line? Is using an LLM to collaboratively build a story alongside the parent and the child bad? I don't know, I suspect not, but that is not a simple question to answer.
It all comes down to whether you see time and attention spent with your kids as a nuisance or the best time you can ever have in your lifetime (I am team B).
And as every child psychologist says, "attention is all you need" :).
Well there is an upper bound of how much time you have available to spend with your child and how you spend that time. You could use an LLM when you can't be there (I suspect this is a bad idea), or you can use it collaboratively with your child (I suspect this is a good idea). I see it as similar to using a phone; your child on wikipedia is a different thing to your child on tiktok.
LLMs are pretty bad at compressing ideas down to the level a smart human can. I think this is quite important for good documentation. Essentially the AI wont do a good job of highlighting the really important and essential details unless you go in and edit the writing by hand afterwards using your human intuition. I am not saying to not use LLMs to write docs, but I would definitely edit to expose the really important details as early as possible in the docs and leave the bulk boring stuff the LLM writes for later.
Yeah, I find it... not incomprehensible, but perhaps deeply misguided the way many people are using LLMs mostly to bulk things up.
1. They generate a lot of fluff that people are accustomed to thinking of as a "sign of work" or "sign of intelligence", but now that counterfeits are easy they are not (should not be) valuable anymore.
2. That fluff (and grammatical annoyances) make readers zone out instead of getting what you want them to get from the document.
3. Much of the document is often an extrapolation from a much smaller set of data like source-code or a spreadsheet or the prompt-stuff. If you really believe the tech will get better, then you should be preserving those higher-truth artifacts instead, and use them to generate an improved extrapolation as-needed.
For sure. But I can definitely produce something good enough to get the job done in a much shorter time by bootstrapping myself from an LLM output and editing from there. (I do edit. Raw LLM output is such a drag to read. Too low signal.)
Luna on xhigh unlocks gh copilot for me. Sol, even discounted, is too expensive to use and is much slower. Luna on the other hand seems to be so cheap you don't have to think about cost.
The problem is it needs world knowledge to know what to lookup. This puts a floor on how little it can know while being able to look up what it doesn't know. Maybe its better if it knows a lot but has a good instinct for verifying that.
How do you actually get to the thing you're looking up? The scrapeable internet is shrinking in response to scrapers.
"Source: rare book ingested and shredded by Anthropic. No, you can't look it up and we can't show you the scan. The remaining open market copy is $5000. Trust me."
reply