ai stack on songwriting: v2 update #2
@@ -14,55 +14,91 @@ Quality and creativity are very subjective properties, and I see complete end-to
|
|||||||
The important thing is I don't and don't plan to use end-to-end (text prompt to audio file) generative AI in what I see as my own recreational creative activities, like songwriting. Instead, I use AI technology in many supportive roles that simply put, makes life easier.
|
The important thing is I don't and don't plan to use end-to-end (text prompt to audio file) generative AI in what I see as my own recreational creative activities, like songwriting. Instead, I use AI technology in many supportive roles that simply put, makes life easier.
|
||||||
I feel it's probably a good idea to disclose how and where I use what sorts of AI. Not that these are classified or confidential or whatever to begin with.
|
I feel it's probably a good idea to disclose how and where I use what sorts of AI. Not that these are classified or confidential or whatever to begin with.
|
||||||
|
|
||||||
**Disclaimer**: this article is not brought to you by any of the projects / services / individuals / organisations mentioned below.
|
**Disclaimer**: This article is not brought to you by any of the projects / services / individuals / organisations mentioned below.
|
||||||
|
***Disclaimer II***: This article is so fucking long you might as well jump to the [quick summary](#Quick_summary) section which is a tl;dr.
|
||||||
|
|
||||||
## Overview
|
## Overview
|
||||||
|
|
||||||
The most controversial, I would imagine, application of modern generative AI in this song is the homebrew English TTS that kind of sounds like Yuzuki Yukari. I never said that is Yukari mind you.
|
The most controversial, I would imagine, application of modern generative AI in this song is the homebrew English TTS that kind of sounds like Yuzuki Yukari. I never said that is Yukari mind you.
|
||||||
The exact repository used for this song (and the previous one, Tunnels) is [this one](https://github.com/RVC-Boss/GPT-SoVITS). I can even specify the commit hash I technically used throughout (haven't really updated since I first cloned) but apparently no one cares about that.
|
The exact repository used for this song (and the previous one, Tunnels) is [this one](https://github.com/RVC-Boss/GPT-SoVITS). I can even specify the commit hash I technically used throughout (haven't really updated since I first cloned) but apparently no one cares about that.
|
||||||
The moral and legal issues around this topic is very messy, but thankfully at least I'm amoral so only the legal concerns are a real thing.
|
The moral and legal issues around this topic are very messy, but thankfully at least I'm amoral so only the legal concerns are real.
|
||||||
|
|
||||||
My day job supposedly is about digital infrastructure maintenance, AI integration into DevOps, and stuff like that. I don't like the job, but I do like the idea of letting machines do the hard labour for me.
|
My day job supposedly is about digital infrastructure maintenance, AI integration into DevOps, and stuff like that. I don't like the job, but I do like the idea of letting machines do the hard labour for me.
|
||||||
Much of my generative AI applications, especially in the example of the very song [pulse under latent semantic envelopes](https://wiki.novoyuuparosk.org/wiki/Pulse_under_latent_semantic_envelopes), are that sort of stuff. To be more exact, I throw specs and reqs at Claude and let it make diagnostic and automation tools for me. At first glance some of them might reek of anything but music.
|
Much of my generative AI applications, especially in the example of the very song [pulse under latent semantic envelopes](https://wiki.novoyuuparosk.org/wiki/Pulse_under_latent_semantic_envelopes), are that sort of stuff. To be more exact, I throw specs and reqs at Claude and let it make diagnostic and automation tools for me. At first glance some of them might reek of anything but music.
|
||||||
|
|
||||||
Another type of help is more directly related to something usually described as songwriting. When I write the lyrics I engage heavily with chatbots. Almost exclusively Claude now (driven by Sonnet 4.6 / Opus 4.7 this time). Obviously I don't tell it to just write everything for me; the finalised version of lyrics is obviously bad like human authored lyrics should be. More details later.
|
Chronologically at the final stages of production my faithful illustrator friend hako told me about [React-Remotion](https://www.remotion.dev/) which is really a thunderstorm that totally changed how I make my music videos. I don't know if I'm going to elaborate on that but it deserves a shoutout for sure. A shoutout to [hako](https://x.com/lxcombox), too. Support your local illustrators who know a thing or two about computers and laser diodes!
|
||||||
|
|
||||||
At the final stages of production my faithful illustrator friend hako told me about [React-Remotion](https://www.remotion.dev/) which is really a thunderstorm that totally changed how I make my music videos. I don't know if I'm going to elaborate on that but it deserves a shoutout for sure. A shoutout to [hako](https://x.com/lxcombox), too. Support your local illustrators who know a thing or two about computers and laser diodes!
|
Aside from the AI-compatible or AI-based tooling, a different type of help is more directly related to something usually described as songwriting. When I write the lyrics I engage heavily with chatbots. Almost exclusively Claude now (driven by Sonnet 4.6 / Opus 4.7 this time). Obviously I don't tell it to just write everything for me; the finalised version of lyrics is obviously bad like human authored lyrics should be. More details later.
|
||||||
|
|
||||||
## Voice synthesis
|
## Voice synthesis
|
||||||
|
|
||||||
|
I know there are people who jump at the honest and objective declaration of 'I used generative AI in creative audio work'. There's no 'hear me out'. More like 'fuck off'.
|
||||||
|
|
||||||
The motive is dead simple: Yukari doesn't have any English TTS package and I so want to write in English and have her do the songs in English. I know VOCALOID 6 is an option but it is not a practical option.
|
The motive is dead simple: Yukari doesn't have any English TTS package and I so want to write in English and have her do the songs in English. I know VOCALOID 6 is an option but it is not a practical option.
|
||||||
|
|
||||||
My requirements for a Yukari English TTS is kind of simple yet niche. The simplicity is that I don't really need a lot of emotional or timbre variations, just the flat Shizuku voice would do; the niche is that I want a very unique accent. Somewhere else I described that as a mix of rallienglanti, Bristolian, and Aussie. In reality what I speak and what I actually want is probably more than that. I speak what I hear (and consider pleasing) and I hear a lot of stuff.
|
My requirements for a Yukari English TTS are kind of simple yet niche. The simplicity is that I don't really need a lot of emotional or timbre variations, just the flat Shizuku voice would do; the niche is that I want a very unique accent. Somewhere else I described that as a mix of rallienglanti, Bristolian, and Aussie. In reality what I speak and what I actually want is probably more than that. I speak what I hear (and consider pleasing) and I hear a lot of stuff.
|
||||||
|
|
||||||
|
By the way, I've been thinking for some time that pursuing the Japanese-style vocal synthesis (TTS) user experience we (the pan-Japanese pan-VOICEROID community) have been so used to, does not translate to a different language. Even not Chinese which is arguably closer to Japanese with the tighter phonology and a relatively overlapping user/fanbase. And spoken English is so far away culturally and parametrically.
|
||||||
|
A.I.Soft did make English and Chinese libs for the Kotonoha sisters, I can only appreciate the effort and will never use them for even the slightest serious En/Cn TTS tasks. That's how awful they are in my opinion. There's a flavour in it but when you don't need the flavour it's just noise.
|
||||||
|
|
||||||
|
### GPT-SoVITS, the theory and the practice
|
||||||
|
|
||||||
Very roughly and factually wrong all over the place, the theory for [GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS) is that the procedure of converting text into speech is further broken down to two steps. **GPT** takes the text and embeds it as pronunciation markers (something like phonemes, but not exactly) and emotions (controls the rises and falls, maybe). **SoVITS** applies the vocal texture based on the intermediate embeddings. The reality is more complicated than that, but the important thing is you can separate accents (handled by GPT) and texture/timbre (handled by SoVITS) into very individual concerns.
|
Very roughly and factually wrong all over the place, the theory for [GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS) is that the procedure of converting text into speech is further broken down to two steps. **GPT** takes the text and embeds it as pronunciation markers (something like phonemes, but not exactly) and emotions (controls the rises and falls, maybe). **SoVITS** applies the vocal texture based on the intermediate embeddings. The reality is more complicated than that, but the important thing is you can separate accents (handled by GPT) and texture/timbre (handled by SoVITS) into very individual concerns.
|
||||||
This setup is what makes possible patching speaker A's voice with speaker B's accents. For the very least you just train the SoVITS weights with A's data, and GPT weights with B's accents. Conceptually concise, innit.
|
This setup is what makes possible patching speaker A's voice with speaker B's accents. For the very least you just train the SoVITS weights with A's data, and GPT weights with B's accents. Conceptually concise, innit.
|
||||||
Theoretically, it is even possible to blend samples from different speakers into one training set. I haven't really done that but I suspect the results would rather be amalgated than homogeneously blended. Apparently cherrypicking samples based on dominant features (e.g. Räikkönen for the tapped 'R's, shouty Aussie for the 'ey-ay' bias, and Mixu Paatelainen for the plosives) could do something good. But I don't know how good that might be and even if it gives me what I have in mind I don't think I have the time (patience) to do the cherrypicking. I just wanted a quickie [sic].
|
|
||||||
|
|
||||||
Disclosure here is that for the voice that dubs the opening two verses and the introduction to signals and noises in the song *pulse*, I used David Cameron Walker's voice for the GPT part. And I haven't bought my Dreamland subscription which I probably should do for my sins.
|
Theoretically, it is even possible to blend samples from different speakers into one training set, but I have not done so.
|
||||||
|
It's also good practice to make sure the training set contains most of the text-to-phoneme mappings you consider important. For example if you want to grasp the tapped Rs in some languages (e.g. Scots, Suomi), you need to have those samples. Or to have a sufficiently big training set like what the megacorps do. Wouldn't say I've done so either.
|
||||||
|
|
||||||
|
For the actual training, I gathered around 5 minutes of speech recordings for each subject (texture and accent) and ran them through the GPT-SoVITS preparation procedures, which is a WebUI and Python backends that helps with noise reduction, normalisation, slicing, ASR tagging, and such.
|
||||||
|
The voice texture donor is my Yukari artifacts for old NCMR songs with spoken Japanese, made with A.I.VOICE. For good precision I made sure I tagged the audio pieces with my original lyrics rather than `whisper` outputs.
|
||||||
|
And the accented English comes from real recordings, which for tagging I heavily relied on whisper ASR but still made crucial corrections where necessary.
|
||||||
|
|
||||||
|
In production, it's as simple as picking the Yukari-trained SoVITS and the Englishman-based GPT models and start generating.
|
||||||
|
Also worth noting that just like many other neural TTS, this GPT-SoVITS package really struggles if the text input is long. Words get skipped or garbled, and pauses get all over the place. My practice is to generate roughly 'one sentence' at a time, and later stitching them all in REAPER. Rhythm correction also happens in REAPER because there's no way to precisely control that in GPT-SoVITS's workspace.
|
||||||
|
|
||||||
|
### On a side note, if you're at all interested
|
||||||
|
|
||||||
|
Disclosure here is that for the voice that dubs the opening two verses and the introduction to signals and noises in the song *pulse*, I used David Cameron Walker's voice to train the GPT model. And I haven't bought my Dreamland subscription, which I probably should do for my sins, when I did the model training.
|
||||||
|
|
||||||
*Anyone that listens to DCW's pods should notice the voice in my song bears almost no similarity to the real Dave Walker, which is kind of a relief to me.*
|
*Anyone that listens to DCW's pods should notice the voice in my song bears almost no similarity to the real Dave Walker, which is kind of a relief to me.*
|
||||||
|
|
||||||
I can at least leave a link to the most dedicated podcast about the (dominantly English) language of football [here](https://www.footballcliches.com/).
|
I can at least leave a link to the most dedicated podcast about the (dominantly English) language of football [here](https://www.footballcliches.com/).
|
||||||
|
|
||||||
***Update:*** *I have subscribed to Dreamland now. For my sins.*
|
***Update:*** *I have subscribed to Dreamland now. For my sins.*
|
||||||
|
|
||||||
## Vibrator coded utilities
|
## Vibrator coded utilities
|
||||||
|
|
||||||
FYI: I don't like the buzzword 'vibe coding'. I think I've only really done it on pure vibes for a very short period. If you produce human-compatible specs almost following the SPICE model of your SWE ones and twos, I think 'vibrator' coding is an understatement.
|
I don't know if indirect generative AI usage during the song development cycle should count or not, but they helped me a lot really. And I guess there will be people who are interested about just generally the environment I have for making music. So here we are.
|
||||||
|
|
||||||
|
### The not-so-generative AI generated by generative AI
|
||||||
|
|
||||||
|
Calling mechanical algorithms and routine scripts 'AI' would seem like an outrageous stretch in today's languages, but that's a poor take from today's languages. They get so associated with the really generative AIs anyway because I guess 90% of the implementations in the past 5 hours of those mechanical algorithms, clusters of ifs and fors, are already products of LLMs.
|
||||||
|
I am not a meteorologist nor a computer scientist, so I do it too. The idea here is that for simple and repetitive tasks like calculating numbers we can use the computer in an old fashioned way with zero LLM involved.
|
||||||
|
|
||||||
|
#### The standalone audio metrics switchblade
|
||||||
|
|
||||||
|
The primary by-product of this song is a metrics toolkit. You actually can find it on [GitHub](https://github.com/mikkelimatlock/uj-mastering-master). This is a Python-Qt tool that analyses several metrics I find useful when mastering / prepping for release a song. Most importantly LUFS.
|
||||||
|
|
||||||
But naming aside, that is basically how I came to finish my long-stalled mastering toolkit, which is the primary by-product of this song. You actually can find it on [GitHub](https://github.com/mikkelimatlock/uj-mastering-master). This is a Python-Qt tool that analyses several metrics I find useful when mastering / prepping for release a song. Most importantly LUFS.
|
|
||||||
The dissatisfaction with Youlean Loudness Meter is that when used in FL Studio it wants me to play over the whole song in 1x speed to calculate the loudness across the whole song. Or I've been doing it wrong. And this is a hassle when the whole song is almost 13 minutes long. So I had to make a tool that spits out results for the 13 minutes in less than 13 minutes, simple as.
|
The dissatisfaction with Youlean Loudness Meter is that when used in FL Studio it wants me to play over the whole song in 1x speed to calculate the loudness across the whole song. Or I've been doing it wrong. And this is a hassle when the whole song is almost 13 minutes long. So I had to make a tool that spits out results for the 13 minutes in less than 13 minutes, simple as.
|
||||||
The basis of the *UJ mastering master* was actually mostly hand-coded by myself ages ago, when it was a very, very crude command line non-interactive script that suddenly pops `matplotlib` plot windows. Peak UX that was, and we have come a long way.
|
The basis of the *UJ mastering master* was actually mostly hand-coded by myself ages ago, when it was a very, very crude command line non-interactive script that suddenly pops `matplotlib` plot windows. Peak UX that was, and I really have come a long way from there.
|
||||||
|
|
||||||
|
LUFS implemented in this utility is mostly BS1770 compliant. I am not a full fan of the official integrated loudness algorithm, but trying to reinvent something that could accommodate any song is not easy.
|
||||||
|
Instead I just kept the 'official' LUFS-I with all the silence discriminations but added a feature where you can just draw your own line and see how loud that reads. Always loved eyeballing metrics.
|
||||||
|
|
||||||
|
#### Automated wiki articles
|
||||||
Another very indirectly associated stack of utilities is my personal Gitea instance that runs on my Pi 5 at home. I felt uneasy about hosting my half-arsed songs on GitHub not because they are a fucking evil corporate, but because my songs are half-arsed. We all need a bit of privacy, no?
|
Another very indirectly associated stack of utilities is my personal Gitea instance that runs on my Pi 5 at home. I felt uneasy about hosting my half-arsed songs on GitHub not because they are a fucking evil corporate, but because my songs are half-arsed. We all need a bit of privacy, no?
|
||||||
*Of course I know there is private repo, I ain't dumb. It's psychologically more assuring to save something like your porn colletion in a physically private location.*
|
*Of course I know there is private repo, I ain't dumb. It's psychologically more assuring to save something like your porn collection in a physically private location.*
|
||||||
The point is I would never get this done properly xor in a way I can remember 30 minutes after I set it all up. So I let Claude Code configurate docker compose yamls and that's mostly it.
|
|
||||||
|
The point is I would never get this done properly xor in a way I can remember 30 minutes after I set it all up. So I let Claude Code configure docker compose yamls and that's mostly it.
|
||||||
I even set up a Gitea action workflow (several workflows with shared bits) for publishing to my wiki from document/writing repos hosted there. That's how you're seeing this page. Or maybe not this exact page.
|
I even set up a Gitea action workflow (several workflows with shared bits) for publishing to my wiki from document/writing repos hosted there. That's how you're seeing this page. Or maybe not this exact page.
|
||||||
|
|
||||||
### A tiny reflection on the whole vibrator thing
|
### A tiny reflection on the whole vibrator thing
|
||||||
|
|
||||||
|
FYI: I don't like the buzzword 'vibe coding'. I think I've only really done it on pure vibes for a very short period. If you produce human-compatible specs almost following the SPICE model of your SWE ones and twos, I think 'vibrator' coding is an understatement. And in all honesty it's not coding anymore. For me it's just 'software development' now.
|
||||||
|
|
||||||
I am always afraid of being rusty of code-writing. But I am equally always anxious about diving too deep into code-writing and failing to see the bigger picture.
|
I am always afraid of being rusty of code-writing. But I am equally always anxious about diving too deep into code-writing and failing to see the bigger picture.
|
||||||
Not one of the first 92 people to say this but leaving the implementations to agents frees me for more tactical and strategical decisions, and that really is a weakness of the agents.
|
Not one of the first 92 people to say this but leaving the implementations to agents frees me for more tactical and strategical decisions, and that really is a weakness of the agents.
|
||||||
Them models and agents are getting better but still, they run into troubles. That is another thing in the agent coding experience that makes me feel I'm still knocking around.
|
Them models and agents are getting better but still, they run into troubles. That is another aspect in this experience that makes me feel I'm still knocking around.
|
||||||
|
|
||||||
## A new age for making glorified slideshows as your music videos
|
## A new age for making glorified slideshows as your music videos
|
||||||
|
|
||||||
@@ -70,13 +106,67 @@ For quite long a time I've been actively (or passively, depends) refusing to mak
|
|||||||
|
|
||||||
But Remotion really is a game changer for me personally. I actually had some experience animating stuff with the HTML/CSS/JS stack. And it's actually amazing how it's almost capable of editing 'real' videos in real time. I haven't really done anything heavy, though, just your usual static assets floating around and dynamically (but deterministically) calculated visual effects.
|
But Remotion really is a game changer for me personally. I actually had some experience animating stuff with the HTML/CSS/JS stack. And it's actually amazing how it's almost capable of editing 'real' videos in real time. I haven't really done anything heavy, though, just your usual static assets floating around and dynamically (but deterministically) calculated visual effects.
|
||||||
|
|
||||||
The point here is that a Remotion 'video' project is described in code, so it's light and agent-friendly. Modern HTML capabilities are mindboggling, and frankly maybe even the Mythos won't be creative an imaginative enough to cook it to its full. So as a human I still need to know the ones and twos of what HTML/CSS is capable of and more importantly what code implementations to call on for the mental images. Or at least try to be fluent in a language so the coding agent gets you.
|
The point here is that a Remotion 'video' project is described in code, so it's light and agent-friendly. Modern HTML capabilities are mindboggling, and frankly maybe even the Mythos won't be creative and imaginative enough to cook it to its full. So as a human I still need to know the ones and twos of what HTML/CSS is capable of and more importantly what code implementations to call on for the mental images. Or at least try to be fluent in a language so the coding agent gets you.
|
||||||
|
|
||||||
## How to have a conversation with an LLM, in the year of 2026, like it was 2019
|
## How to have a conversation with an LLM, in the year of 2026, like it was 2019
|
||||||
|
|
||||||
I consider myself open to new technologies and ideas yet slow to make the initial adaptation. My reasoning is that some of those on the technical level are no more than makeshifts that won't be required that much once a paradigm shift or significant breakthrough is there one a higher level.
|
I consider myself open to new technologies and ideas yet slow to make the initial adaptation. My reasoning is that some of those on the technical level are no more than makeshifts that won't be required that much once a paradigm shift or significant breakthrough is there on a higher level.
|
||||||
An example would be the so-called prompt engineering. To hell with that is all I'll say.
|
An example would be the so-called prompt engineering. To hell with that is all I'll say.
|
||||||
|
|
||||||
I'm still not fully embracing clustered multi-agent tasking, skills-based enforcing era of agentic working. Nor that my personal uses really need that much overhead. These might happen if I ever switch to a self-hosted agent framework like Hermes, but Claude Code, in honesty, does most of the memory/context stuff for me.
|
I'm still not fully embracing clustered multi-agent tasking, skills-based enforcing era of agentic working. Nor that my personal uses really need that much overhead. These might happen if I ever switch to a self-hosted agent framework like Hermes, but Claude Code, in honesty, does most of the memory/context stuff for me.
|
||||||
Songwriting, the literal part, is even less suited to being made into a model, a set of procedures, a paradigm, or anything rigid. For really localised sub-tasks like typesetting or grammar checks it's probably OK to use a routine skill, but for the softer, more creative parts, I believe simple and natural conversations is the way to go.
|
Songwriting, the literal part, is even less suited to being made into a model, a set of procedures, a paradigm, or anything rigid. For really localised sub-tasks like typesetting or grammar checks it's probably OK to use a routine skill, but for the softer, more creative parts, I believe simple and natural conversations is the way to go.
|
||||||
That, and actively instructing about the inner mechanism of an LLM agent rather than truly simple and natural conversations. It's just in the blood at this point now.
|
That, and actively instructing about the inner mechanism of an LLM agent rather than truly simple and natural conversations. It's just in the blood at this point now.
|
||||||
|
|
||||||
|
### So, how to have that conversation?
|
||||||
|
|
||||||
|
The principle: you have to write yourself and ask for feedback, suggestions, reviews... responsive stuff that still won't make it into the final product. Creative writing is not regulated, and I think it's something the LLMs still struggle at.
|
||||||
|
|
||||||
|
My personal take is that an author bases all the literal (and musical) building blocks, the vocabulary, the composition, the imagery, etc., on a sentimental and semantic basis. When you try to convert that cluster of emotions and feelings and thoughts into language, it is always lossy. Speaking buzzwords, the **vibe** is a priori, and the textualisation of it is a posteriori.
|
||||||
|
'Human-in-the-loop' was a hot concept, maybe not so much nowadays, but it is so true for computer assisted creative writing. The nuance here is perhaps if you label it as only an assistance, something needs to take the helm, to make out what that vibe, the target, actually is. The computer is just there to take over those responsibilities of the flesh database, the library visits, all that physical and data link layer actions.
|
||||||
|
|
||||||
|
#### Do not throw concrete blocks of thoughts from the top of your house, to the LLMs
|
||||||
|
|
||||||
|
Before we fall into the philosophical rabbit hole of whether LLM instances are sentient individuals and have individual minds, let's pretend that they are similar to someone who simply is not the author. When approaching an arbitrary piece of text, this second individual can only attack from the surface and try to work out what's between and behind the lines.
|
||||||
|
|
||||||
|
The tricky thing here is that, aside from the apparent formal lyrics, the conversational language that is supposed to convey that blob of thoughts, the vibe, is also text. It's a best effort situation trying to transport the authentic underlying stuff with text. I personally think that's why so many philosophers resorted to meticulous and mind paralysing word play. It kind of rose as the only way to shove one centimetre deeper into human minds.
|
||||||
|
And it's often the case that said target is not clear at the first few days of a project. The tendency for LLM (agents) of being overly eager to lock in onto some very specific keywords is actually helpful to me. That behaviour helps me steer away from overfitting my fuzzy blob of thoughts onto a strangely specific curvature of concrete concepts. Take the *pulse* song itself as an example, Claude Opus 4.7 was so fucking opinionated that what I was trying to describe was definitely a submarine, and threw at me all your periscopes and ballasts and Walter's hydrogen peroxide drive systems<b>\*</b>.
|
||||||
|
|
||||||
|
*<b>\*</b>: Claude didn't pop that up, but I read [a great blog about it on wwiiafterwwii](https://wwiiafterwwii.wordpress.com/2026/03/21/walter-u-boat-technology-after-wwii/). Had to sneak it in.*
|
||||||
|
|
||||||
|
If you're not a steadfast writer that can detect 'no, it's not right' at first or second notice, that could be rather problematic. If you are, it's problematic in a different way because the LLM can't talk about a cloud of bodiless thoughts and will always try to condense that into something more specific.
|
||||||
|
I actually gave up offloading the task for keeping the 'fuzzily bounded region of thoughts' in its shapeless shape to any LLM or even second human individual. It's mostly a personal problem of not being able to convert it into words in an efficient way, but it's a lingering problem and presumably unsolvable. It's my thoughts only when it's in my mind and I best keep it there.
|
||||||
|
|
||||||
|
#### But save yourself all the time for the mechanical works
|
||||||
|
|
||||||
|
I already talked about how to milk time for an additional cuppa tea by letting Claude Code do the utility wiring and plumbing. The same spirit can be translated to literal chores like picking the right words, grammar checks, rhythmic and rhyme matching, albeit to a lower level of automation. I had to keep the conversation going because this one is heavily Human-in-the-loop.
|
||||||
|
|
||||||
|
It's difficult to generalise so I'll just explain with an example. Take this couplet from our new favourite 12'53" song *pulse under latent semantic envelopes*:
|
||||||
|
```
|
||||||
|
Lossy lies lament the ludicrous lawful liberty
|
||||||
|
Siren sounds soaken in the sauna of substantiality
|
||||||
|
```
|
||||||
|
This is part of the acrostic verse which spells 'PULSE'. And for no specific reason I just decided that this two-liner will be alliterations with a bit of freedom for cements.
|
||||||
|
|
||||||
|
I basically started with:
|
||||||
|
```
|
||||||
|
Lossy l.. lament the ludicrous lawful l..
|
||||||
|
Siren sounds sunken ... sauna something
|
||||||
|
```
|
||||||
|
S is a good letter because it lets me sneak in the sauna.
|
||||||
|
Then I told Claude Code that:
|
||||||
|
> - The two lines will be alliterations but connecting words are exempted.
|
||||||
|
> - `ludicrous` and `sauna` are must-keeps.
|
||||||
|
> - Should serve as an emotional transition between the two lines above and the closing line. (at that time not fully confirmed, but the embedding was already there)
|
||||||
|
>
|
||||||
|
> So you should give me some words starting with `L` and some more starting with `S`. Follow the generic direction in the draft.
|
||||||
|
|
||||||
|
Well, roughly that, I can't replicate my exact prompts. But the idea here is to hold the reins of expression and control where it heads. Doesn't have to be exact if you're still not sure, but OK to pin down as many constraints you feel like keeping.
|
||||||
|
The problem is too many bobby pins will result in either a mission impossible or a complete sideways result, which is the only sane answer to an unanswerable question. Imagine requesting a zoom lens system that does everything from 8mm to 450mm, within an overall length of 450 mm, and a max aperture of F/1.7. That is nonsense and optical software will tell you so. The LLMs usually don't but give you shite you most probably don't want.
|
||||||
|
And that's where the author must make difficult decisions, or display his superhuman flexes with words. Common people like me just make sacrifices and reduce the 'must fulfill' list to a reasonable version. Like I must have the 'L' and 'S' alliterations, but with a bit of freedom; I need to keep the 'ludicrous' and 'sauna' words specifically; the general emotion should transition not too abruptly (and that's one for me to judge, not entirely on Claude).
|
||||||
|
|
||||||
|
Some people like to just give Mr. AI a very, very concise prompt and take whatever he spits out, I don't do that; some other people try to generalise the process and save them as skill MD files, I don't do that either. My practice should be described as the fine-grained level of control I feel most comfortable with and nothing more than that.
|
||||||
|
|
||||||
|
## Quick summary
|
||||||
|
|
||||||
|
Creativity is still a largely human attribute, but it's safe and helpful to offload the tooling and objective aspects of creating something to AIs. Generative or not.
|
||||||
|
Copyright is a rather vague topic to code, but my stance is you shouldn't let constructed morals get in the way of what you see crucial.
|
||||||
Reference in New Issue
Block a user