Comment by hbn

Comment by hbn 4 days ago

119 replies

I've enjoyed using ffmpeg 1000% more since I was able to stop doing manually the tedious task of Googling for Stack Overflow answers and cobbling them into a command and got Chat GPT to write me commands instead.

simonw 4 days ago

I use ffmpeg multiple times a week thanks to LLMs. It's my top use-case for my "llm cmd" tool:

  uv tool install llm
  llm install llm-cmd

  llm cmd use ffmpeg to extract audio from myfile.mov and save that as mp3
https://github.com/simonw/llm-cmd
  • resonious 4 days ago

    I tried this (though with a different tool called aichat) for extremely simple stuff like just "convert this mov to mp4" and it generated overly complex commands that failed due to missing libraries. When I removed the "crap" from the commands, they worked.

    So much like code assistance, they still need a fair amount of baby sitting. A good boost for experienced operators but might suck for beginners.

    • cm2187 3 days ago

      Plus you need to know the format of your source file to design the command correctly. How many audio tracks, is the first video track a thumbnail or the video, are the subtitles tracks forced, etc.

      And in some situations ffmpeg has some warts you have to go around. Like they introduced recently a moronic change of behaviour where the first sub tracks becomes forced/default irrespective of the original forced/default flag of the source. You need to add "-default_mode infer_no_subs" to counter that.

      • Philpax 3 days ago

        I usually just paste the output of `ffprobe` into Claude when it's ambiguous. Works a treat.

    • Over2Chars 4 days ago

      My feelings exactly, but I think that's OK!

      It's another tool and one that might actually improve with time. I don't see GNU's man pages getting any better spontaneously.

      Whoa, what if they started to use AI to auto-generate man pages...

      • BlaDeKke 3 days ago

        > Whoa, what if they started to use AI to auto-generate man pages...

        That’s the time to start my career in woodworking.

      • michaelcampbell 2 days ago

        > what if they started to use AI to auto-generate man pages...

        Then they'd be wrong about 20% of the time, and still no one would read them. ;-)

        (NB: I'm of the age that I do read them, but I think I'm in the minority.)

    • BiteCode_dev 3 days ago

      Reading this feels like seing a guy getting his first car in 1920 and complaining he still has to drive it himself.

      • imiric 3 days ago

        To me it's more like a guy getting his first car and complaining that the car is driving him in a direction that may or may not be correct, despite his best efforts to steer it where he wants to go. And the only way to know whether he ends up in the right place is to get out of the car, look around, and maybe ask more experienced drivers. Failing that, his only option is to get back in and hope to be luckier in the next trip.

        Or he can just ditch the car and walk. Sure, it's slower and requires more effort, but he knows exactly how to do that and where it will take him.

      • LeoWattenberg 3 days ago

        The beer brewers in my home town used to have a self-driving horse and cart which knew the daily delivery route going by all pubs and didn't really need a human to steer it or indeed be conscious during the trip. Expectedly, the delivery guy would get drunk first thing in the morning and just get carted about collecting the money.

      • pbhjpbhj 3 days ago

        Pony & trap could be largely self-driving, after an initial training period. That would have been a distinct negative to "upgrading" for some, I'd imagine.

        • nine_k 3 days ago

          It's speed and load capacity vs self-driving.

          If we could imagine wiring a pony to control a car, its brain, while good at navigation, would likely be inadequate at the speed that a car attains.

      • atoav 3 days ago

        Sell that guy probably got carried home by his horse after drinking half a bottle of whiskey, so maybe he had a point.

      • [removed] 2 days ago
        [deleted]
      • shriek 3 days ago

        Or maybe calling a cab and telling the cab driver each direction to get to the destination instead of the cab driver just taking you there.

    • assimpleaspossi 3 days ago

      My experience exactly.

      I no longer check with these AI tools after a number of attempts. Unrelated, a friend thought there was a NFL football game last Saturday at noon. Checking with Google's Gemini, it said "no", but there was one between two teams whose season had ended two weeks before at 1:00 Eastern Time and 2:00 Central. (The times are backwards.)

      • bityard 3 days ago

        Do LLMs have knowledge of current events?

    • sdesol 4 days ago

      > "convert this mov to mp4"

      Did any of the commands look like the ones in the left window:

      https://beta.gitsense.com/?chats=12850fe4-ffb1-4618-9215-c13...

      The left window contains a summary of all the LLMs asked, including all commands. The right window contains the individual LLM responses.

      I asked about gotchas with missing libraries as well, and Sonnet 3.5 said there were. Were these the same libraries that were missing for you?

      • resonious 4 days ago

        Looking at this, I am pretty sure I also received a "libx264" clause. Removing it made the command work for me.

    • keeganpoppen 3 days ago

      what exactly do you want the llm to do here? if the ask was so unambiguous and simple that it could be reliably generated, then the interface wouldn't be so complicated to use in the first place! LLMs are not in any way best suited for one-shot prompt => perfect output, and expectations to that effect are extremely unreasonable. the reason why LLMs are still hard for beginners to use is because the software is hard to use correctly. as with LLM output goes life itself: the results you get from using a tool can only ever be as good as the (mental) model used to choose that tool & the inputs to begin with. if all the information required to generate the output were contained by the initial prompt, then there would be absolutely no need to use the LLM at all in the first place.

    • Philpax 4 days ago

      Hate to be that guy, but which LLM was doing the generation? GPT-4 Turbo / Claude 3.x have not really let me down in generating ffmpeg commands - especially for basic requests - with most of their failures resulting from domain-specific vagaries that an expert would need to weigh in on m

      • th0ma5 4 days ago

        Hate to be that guy, but which model works without fail for any task that ffmpeg can do?

  • pmarreck 3 days ago

    A while back I simply wrote my own bash function for this called `please`

    as in

        bash> please "use ffmpeg to extract audio from myfile.mov and save it as mp3"
    
    It will then courteously show you the command it wants to run before you agree to do it.

    Here is the whole thing, with its two dependent functions, so that people stop writing their own versions of this lol. All it needs is an OPENAI_API_KEY, feel free to modify for other LLMs

    EDIT: Moved to a gist: https://gist.github.com/pmarreck/9ce17f7996347dd532f3e20a2a3...

    Suggestions welcome- for example I want to add a feature that either just copies it (for further modification) or prepopulates the command line with it somehow (possibly for further modification, or even for skipping the approval step)

    • smusamashah 3 days ago

      please is such an appropriate name. Will rename my ChatGPT alias to please.

  • atoav 3 days ago

    Did you just invent the LLM-equivalent of curl-piping unread shell scripts into sh?

    I am sure that will never cause any problems.

    • bspammer 3 days ago

      It displays the generated command to you, there's an additional step to confirm.

    • ykonstant 3 days ago

      > Did you just invent the LLM-equivalent of curl-piping unread shell scripts into sh?

      Many such cases.

  • dekhn 4 days ago

    "The future is already here. It's just not very well distributed"

    (honestly, the work you share is very inspiring)

  • zahlman 4 days ago

    >This will then be displayed in your terminal ready for you to edit it, or hit <enter> to execute the prompt. If the command doesnt't look right, hit Ctrl+C to cancel.

    I appreciate the UI choice here. I have yet to do anything with AI (consciously and deliberately, anyway) but this sort of thing is exactly what I imagine as a proper use case.

    • hnuser123456 4 days ago

      Just like all other code. There will be user-respecting open source code and tools, and there's user-disrespecting profitable closed code that makes too many decisions for you.

  • th0ma5 4 days ago

    You should figure out what went wrong for the other commenter and fix your tool.

  • mrweasel 3 days ago

    While I love that that works, I still feel like just maybe ffmpeg needs a better interface. Not necessarily a GUI, just a better designed command line.

  • Waterluvian 4 days ago

    I think I’m finally sold on actually attempting to add some LLM to my toolbelt.

    As a helper and not a replacement, this sounds grand. Like the most epic autocomplete. Because I hate how much time I waste trying to figure out the command line incantation when I already know precisely what I want to do. It’s the weakest part of the command line experience.

    • Over2Chars 4 days ago

      But possibly the most rewarding. The struggle is its own reward that pays off later many times over.

      • hbn 3 days ago

        There are times I feel minor guilt for using an LLM to relieve brainwork, like figuring out an algorithm. That's probably a skill I should continue practicing for my own sake.

        ffmpeg commands though? It's really not a practical skill outside of using ffmpeg. There's nothing really rewarding to me about memorizing awkwardly designed CLI incantations. It's all arbitrary.

        • Over2Chars 3 days ago

          memorization is not what I'm talking about.

          I'm talking about

          1. you have a problem, you try something and it doesn't work. 2. you find an LLM and it "gives" you the answer with one or two tries 3. problem solved! what have you learned? How to have answers given to you when you ask.

          or 2. you look for an answer in a dizzying haze of man pages, quacks, and website Q&As. 3. you try and try again and eventually problem solved. you have learned not only how to solve a particular problem but your overall ability to solve similar problems has done up. You've learned how to fish, not just ask an LLM for a fish.

      • Waterluvian 4 days ago

        Not for me. It’s a tool I don’t care to use any more than I have to. I’m much more interested in what I’m using the tool to accomplish.

levocardia 4 days ago

For the longest time I had ffmpeg in the same bucket as regex: "God I really need to learn this but I'm going to hate it so much." Then ChatGPT came along and solved both problems!

  • zxvkhkxvdvbdxz 4 days ago

    Interesting. Being able to use regexps for text processing through my career has probably saved me a few thousand hours of programming one-off solutions so far. It is one of those skills that really pays off to learn proper.

    And speaking of ffmpeg, or tooling in general, I tend to make notes. After a while you end up with a pretty decent curated reference.

    • codetrotter 4 days ago

      I use regexes a lot. The main thing that always trips me up is dealing with escaping, because different tools I use – vim, sed, rg, and so on – sometimes have different meanings for when to escape or not.

      In one tool you’ll use + to match one or more times, and \+ to mean literal plus sign.

      In another tool you’ll use \+ to match one or more time, and + to mean literal plus sign.

      In one tool you’ll use ( and ) to create a match group, and \( and \) to mean literal open and close parentheses.

      In another tool you’ll use \( and \) to create a match group, and ( and ) to mean literal open and close parentheses.

      This is basically the only problem I have when writing regexes, for the kinds of regexes I write.

      Also, one thing that’s not a problem per se but something that leads me to write my regexes with more characters than strictly necessary is that I rarely use shorthand for groups of characters. For example the tool might have a shorthand for digit but I always write [0-9] when I need to match a digit. Also probably because the shorthand might or might not be different for different tools.

      Regexes are also known to be “write once read never”, in that writing a regex is relatively easy, but revisiting a semi-complicated regex you or someone else wrote in the past takes a little bit of extra effort to figure out what it’s matching and what edits one should make to it. In this case, tools like https://regex101.com/ or https://www.debuggex.com/ help a lot.

      • nuancebydefault 3 days ago

        The problem with escaping (like with using quotes) is often that you need to know through how many parsers the string goes. The shell or editor, the language you are programming in and the regexp engine each time can strip off an escape character or a set of outer quotes. That and of course different dialects of regexp makes things complicated.

      • account42 2 days ago

        This is like saying that words have different meaning if you talk to someone in english or french. Most tools have a switch to inform them that you're going to talk in english (perl-compatible regular expressions).

    • mystified5016 4 days ago

      No one doubts the power or utility of regexes or ffmpeg, but they are both complicated beasts that really take a lot of skill.

      They're both tools where if they're part of your daily workflow you'll get immense value out of learning them thoroughly. If instead you need a regex once or twice a week, the benefit is not greater than the cost of learning to do it myself. I have a hundred other equally complicated things to learn and remember, half the job of the computer is to know things I can't put in my brain. If it can do the regex for me, I suddenly get 70% of the value at no cost.

      Regex is not a tool I need often enough to justify the hours and brain space. But it is still an indespensible tool. So when I need a regex, I either ask a human wizard I know, or now I ask my computer directly.

      • imp0cat 3 days ago

        It's self-reinforcing though. If you invest the time to learn, then you may find yourself (i a beafutiful house :)) using it a lot more than two times a week.

      • [removed] 3 days ago
        [deleted]
  • earnestinger 4 days ago

    Not sure about ffmpeg, but you should definitely try memorising regexp. Casual Search&replace that becomes possible is worth it.

    • sergiotapia 4 days ago

      in 15 years it never sticks and by the time i need it again i've forgotten it! :D

      • kevin_thibedeau 4 days ago

        Don't learn the Perl influenced extensions. You just need POSIX EREs (and BREs for some older utilities) which are simple enough to keep in the head.

  • teaearlgraycold 4 days ago

    Gotta be honest, years of configuring automod on Reddit have honed me into a regex God.

  • jmb99 3 days ago

    For me, it wasn’t so much learning ffmpeg, as it was understanding containers/codecs/encoders/streams/etc. Learning all of the intricacies there made ffmpeg make a lot more sense.

    • skydhash 3 days ago

      Almost no one cares to understand the domain of the tool anymore, they only want result and expect a simplified interface that already does the unique thing they want to do, but can’t accept that a power tool can only be used with training.

  • hackingonempty 4 days ago

    CSS has entered the ChatGPT.

    • kccqzy 4 days ago

      My rule for using LLMs is that anything that's one off is okay. Anything that's more permanent and committed to a repo needs a human review. I strongly suggest you have an understanding of the basics (at least the box model) so that you are competent at reviewing CSS code before using LLM for that.

    • permo-w 4 days ago

      I've been looking for a good guide on prompting LLMs for CSS.

      does anyone know of any?

      • johnisgood 3 days ago

        I have no set of rules when prompting LLMs for CSS, it does seem to work more or less for me though.

        What are your current issues or what limitations have you ran into?

        • permo-w 3 days ago

          mainly with trying to impart a particular art style.

juancroldan 4 days ago

Same here, it's one of these things where AI has taken over completely and I'm just a broker that copy-pastes error traces.

magarnicle 4 days ago

My experience got even better once I learned how complex filters worked.

  • dylan604 4 days ago

    learning how to use splits to do multiple things all in one command is a god send. the savings of only needed to read the source and convert to baseband video once is a great savings.

    i started with avisynth, and it took time for my brain to switch to ffmpeg. i don't know how i could function without ffmpeg at this point

NetOpWibby 3 days ago

Truly, a net positive to my life. Just a few days ago I asked my AI buddy (Claude) to create a zsh script to organize my downloads folder according to the Johnny Decimal system. I’ve since modified it to move the files to a JD setup on my desktop.

The sense of elation I get when I wonder aloud to my digital friend and they generate what I thought was too much to expect. Well worth the subscription.

Over2Chars 4 days ago

I think you're onto something. I've had hit or miss experiences with code from LLMs but it definitely makes the searching part different.

I had a problem I'd been thinking about for some time and I thought "Ill have some LLM give me an answer" and it did - it was wrong and didn't work but it got me to thinking about the problem in a slightly different way and my quacks after that got me an exact solution to this problem.

So I'm willing to give the AI more than partial credit.

bambax 3 days ago

Basic syntax for re-encoding a video file did take me some time to memorize, but isn't in fact too hard:

  ffmpeg <Input file(s)> <Codec(s)> <MAPping of streams> <Video Filters> output_file
- input file: -i, can be repeated for multiple input files, like so:

  ffmpeg -i file1.mp4 -i file2.mkv
If there is more than one input file then some mapping is needed to decide what goes out in the output file.

- codec: -c:x where x is the type of codec (v: video, a: audio or s:subtitles), followed by its name, like so:

  -c:v libx265
I usually never set the audio codec as the guesses made by ffmpeg, based on output file type, are always right (in my experience), but deciding the video codec is useful, and so is the subtitles codec, as not all containers (file formats) support all codecs; mkv is the most flexible for subtitles codecs.

- mapping of streams: -map <input_file>:<stream_type>:<order>, like so:

  -map 0:v:0 -map 1:a:1 -map 1:a:0 -map 1:s:4
Map tells ffmpeg what stream from the input files to put in the output file. The first number is the position of the input file in the command, so if we're following the same example as above, '0' would be 'file1.mp4' and '1' would be 'file2.mkv'. The parameter in the middle is the stream type (v for video, a for audio, s for subtitles). The last number is the position of the stream IN THE INPUT FILE (NOT in the output file).

The position of the stream in the output file is determined by the position of the map command in the command line, so for example in the command above we are inverting the position of the audio streams (taken from 'file2.mkv'), as audio stream 1 will be in first position in the output file, and audio stream 0 (the first in the second input file) will be in second position in the output file.

This map thing is for me the most counter-intuitive because it's unusual for a CLI to be order-dependent. But, well, it is.

- video filters: -vf

Video filters can be extremely complex and I don't pretend to know how to use them by heart. But one simple video filter that I use often is 'scale', for resizing a video:

  -vf scale=<width>:<height>
width and height can be exact values in pixels, or one of them can be '-1' and then ffmpeg computes it based on the current aspect ratio and the other provided value, like this for example:

  -vf scale=320:-1
This doesn't always work because the computed value should be an even integer; if it's not, ffmpeg will raise an error and tell you why; then you can replace the -1 with the nearest even integer (I wonder why it can't do that by itself, but apparently, it can't).

And that's about it! ffmpeg options are immense, but this gets me through 90% of my video encoding needs, without looking at a manual or ask an LLM. (The only other options I use often are -ss and -t for start time and duration, to time-crop a video.)

  • izacus 3 days ago

    > This doesn't always work because the computed value should be an even integer; if it's not, ffmpeg will raise an error and tell you why; then you can replace the -1 with the nearest even integer (I wonder why it can't do that by itself, but apparently, it can't).

    It's not about integer, but some of the sizes need to be even. You can use `-vf scale=320:-2` to ensure that.

  • jmb99 3 days ago

    > then you can replace the -1 with the nearest even integer (I wonder why it can't do that by itself, but apparently, it can't).

    Likely because the aspect ratio will no longer be the same. There will either be lost information (cropping), compression/stretching, or black bars, none of which should be default behaviour. Hence, the warning.

sathishvj 3 days ago

I would like to throw in a tool that I built into the ring: gencmd - https://gencmd.com/. There is a web version and also a CLI version.

If the CLI is installed, you can do: gencmd -c ffmpeg extract first 1 minute of video

Or you can just search for the same in the browser page.

[removed] 4 days ago
[deleted]
nine_k 3 days ago

I do it the old way: I write down the commands as a shell script, and reuse later.

But really what ffmpeg is missing is an expressive language to describe its operation. Something well-structured, like what jq does for JSON.

  • skydhash 3 days ago

    It already does. It’s the cli flags. What you’re missing is the semantic which you can get with learning about containers, codecs, and other stuff. You don’t use grep and sed with no understanding of what a text file is.

michaelcampbell 2 days ago

ffmpeg and jq are 2 commands I've about given up trying to "use" with any facility and am more than happy to pawn that off to one of the Gippity's; chat, claude, etc.

urda 4 days ago

For me it was using a container of it, instead of having to install all the things FFmpeg needs on a machine.

archerx 3 days ago

Why not just use Handbrake? It’s just FFMpeg but with a GUI.

  • wildzzz 3 days ago

    That's fine for encoding but Handbrake doesn't let you do video streaming to my knowledge.

skirge 4 days ago

llm - Clippit of 202x, but for the original Pentium was enough.