top of page

Generally AI – Season 2, Episode 1: Generative AI and Creativity

The second season of Generally AI opens with a question that now reaches far beyond research labs: What happens to creativity when software can produce books, songs, samples, essays, and images almost instantly? The hosts approach the subject through practical experiments, technical evaluation methods, and examples of synthetic media already competing for attention online.

Their conversation reveals two different stories unfolding at once. Generative AI can flood marketplaces with disposable material, but it can also give artists an unusually flexible source of raw ideas. Understanding that distinction requires looking beyond whether a machine can generate something plausible and asking whether its output is useful, controllable, original, and worth a human audience’s time.

AI-Generated Books and the Economics of Abundance

The episode begins with a Dutch investigation into books sold online. Of roughly 323,000 titles examined, about 2% appeared to be AI-generated. The estimated share had risen sharply: from approximately 0.1% before ChatGPT’s arrival to 4.2% in April 2024.

Those numbers illustrate how quickly publishing changes when producing a manuscript no longer demands months of work. The hosts point to Amazon’s limit of three books per author per day. That restriction may curb industrial-scale uploads, yet it remains an extraordinary production rate by ordinary writing standards.

Quality is another concern. Many of the suspected AI books received ratings of two stars or fewer. In one especially revealing mistake, a customer reportedly received a book containing an AI instruction rather than the intended prose—evidence that someone had published the material without even reviewing the final text.

The problem is not merely an increase in mediocre books. Cheap synthetic titles can crowd search results and make credible work harder to find. Near the end of the episode, the hosts describe searching a Dutch bookseller for an Oppenheimer biography and encountering AI-made alternatives instead of prominently seeing American Prometheus, the acclaimed biography associated with the film. They are particularly troubled that health, self-help, management, and personal development are among the categories attracting AI-generated publications. Errors in those fields can cause more harm than an uninspired novel.

From Novelty Songs to an Endless Stream of Synthetic Music

AI music has advanced considerably since the podcast’s first season. The hosts discuss Suno and play with the idea of generating catchy, polished songs around almost any prompt. A cat-themed track demonstrates the technology’s immediate appeal: it is upbeat, memorable, and easy to enjoy without knowing how it was assembled.

That accessibility also enables extreme output. The episode cites a creator releasing multiple cat-themed albums on Spotify in a single day. If such practices become common, streaming services may eventually need upload restrictions comparable to those introduced by book platforms.

One host experiments with treating generated songs as a substitute for conventional streaming music. The experience is entertaining at first, but the relentless pop sensibility soon becomes tiring. This points to a limitation hidden by the initial surprise: music can sound competent while still lacking range, restraint, and a durable artistic identity.

The distinction matters. A generator capable of producing a pleasant song is not necessarily creating a body of work people will revisit. Novelty earns attention once; meaningful variation and artistic intention sustain it.

AI Works Best as Creative Material, Not a Finished Product

The episode becomes more optimistic when it turns from automatic song production to tools that musicians can actively manipulate. The hosts discuss a Google demonstration in which AI-generated musical elements were mixed with live loops and performance. Here, the machine functions less like an autonomous composer and more like an expandable instrument.

That model of collaboration offers a stronger creative proposition. A musician might generate an unfamiliar texture, isolate a fragment, change its rhythm, combine it with an existing loop, and transform it into something the original system could not have anticipated. The human remains responsible for selection, structure, taste, and direction.

AI sample generators make this workflow tangible. One service discussed in the episode produces short audio clips from descriptions such as a funky melody, a piano phrase, or a particular string texture. The resulting two-to-four-second sounds can be loaded into a sampler and rearranged into beats.

The hosts compare the process to searching through infinitely many crates of records. Unlike traditional crate digging, however, the source material can be requested rather than discovered. A Teenage Engineering PO-33 sampler provides a hands-on demonstration of how imperfect generated clips can become ingredients in a larger composition.

This is a more productive definition of AI-assisted creativity: not pressing a button for a finished artifact, but using generation to enlarge the space of materials available for human decisions.

Why Audio Generation Still Resists Precise Direction

The sample experiments also expose the technology’s weaknesses. Outputs are inconsistent, and producing a usable sound may require repeated attempts. A clip can begin too early, arrive with an awkward buildup, or contain more activity than requested.

The hosts describe the difficulty of asking for a single clean sound. A prompt for one duck quack, for example, may yield several quacks or a noisy sequence. The model seems eager to fill the available duration instead of respecting the musical value of silence.

Some artifacts may arise during the process of converting model outputs into an audio waveform. Token-based generation and decoding can introduce noise or reduce fidelity. Yet the larger obstacle is interaction. A human collaborator can understand requests such as “make the bass darker,” “hold that note longer,” or “leave more space before the hit.” Current tools do not consistently support that kind of iterative, musically aware feedback.

The hosts imagine systems that could accompany a bass line with drums or strings, perhaps even inside an effects pedal. Such a tool would allow musicians to jam with a responsive machine rather than repeatedly submit isolated text prompts. The episode suggests that this real-time partnership remains more aspiration than dependable product.

Text-to-Sample Tools and the Copyright Question

Another text-to-sample application discussed in the episode can run as a standalone program or plug-in, allowing generated clips to be dragged into virtual instruments and drum kits. A locally running version downloads a music-generation model to a MacBook, showing how permissively licensed models can become components inside many different production tools.

Generated samples may also offer an alternative to extracting recordings from existing songs. Traditional sampling can require complicated rights clearance, especially for musicians whose work depends heavily on borrowed fragments. Albums such as the Beastie Boys’ Paul’s Boutique demonstrate how creatively powerful sampling can be, while also recalling an era whose licensing environment was very different.

Synthetic samples do not automatically eliminate every legal or ethical question. Nevertheless, generating new source audio could reduce reliance on recognizable copyrighted recordings and give sample-based artists more material to reshape.

Measuring Creativity Without a Single Correct Answer

The episode’s second major theme is evaluation. A conventional classifier can be tested against known labels, and a regression model can be compared with expected numerical values. Creative generation rarely offers such a clear target. There is no universally correct song, essay, or image for a given prompt.

Language-model training still provides measurable signals. Next-token prediction produces a loss value, and scaling laws relate expected performance to model size, training data, and computation. But lower loss is only a proxy. Users ultimately care whether a model answers questions, follows instructions, summarizes accurately, or creates something worthwhile.

Benchmarks such as MMLU test knowledge through questions with checkable answers, while public leaderboards compare models across multiple tasks. These methods work best when success can be objectively scored. They are less decisive for open-ended writing.

Text metrics offer partial solutions. BLEU and ROUGE compare generated text with references through overlapping words or word sequences, emphasizing different aspects of precision and recall. BERT-based scoring instead compares semantic representations, making it possible to recognize similar meanings expressed in different language. Even so, resemblance to a reference does not necessarily equal quality or creativity.

Image generation faces the same dilemma. Fréchet Inception Distance compares statistical representations of generated and reference-image collections. CLIP-based scoring estimates whether an image corresponds to its prompt. Both are useful, but neither fully captures visual coherence, originality, or human preference. A model used as a judge also inherits its own weaknesses—for example, difficulty counting objects accurately.

Human Preference, Model Rankings, and RLHF

When objective scoring reaches its limits, the hosts return to human judgment. People can compare two outputs side by side and choose the better one—the episode’s “optometrist” approach. Repeated comparisons can produce an Elo-style ranking even when no absolute quality score exists.

The same principle supports reinforcement learning from human feedback. Evaluators rank several responses to a prompt, creating preference data that can be used to make a language model more helpful and better aligned with user intent.

Researchers can also ask one model to evaluate another, potentially reducing the cost of human review. But this introduces a circular problem: a machine judge may reproduce biases, overlook errors, or reward the same superficial qualities as the model being tested. Automated evaluation can scale oversight, but it does not make human standards unnecessary.

The hosts’ practical conclusion is that AI succeeds when it helps people accomplish what they intended. For creative systems, the final test is therefore not merely technical plausibility. It is whether people find the result valuable.

The Clever Hans Warning for Generative AI

A closing discussion about animal intelligence supplies an apt caution. Clever Hans, a horse once believed capable of arithmetic, was eventually shown to be responding to subtle cues from nearby humans. Observers had attributed sophisticated reasoning to behavior produced through a different mechanism.

Generative AI invites a similar interpretive risk. Fluent language or attractive media can lead people to project intention and understanding onto the system. Sometimes the most creative act occurs in the human observer, who supplies meaning to an ambiguous output.

That does not make the technology useless. It clarifies where its value lies. Generative systems can produce an astonishing quantity of material, but people still determine which fragments deserve attention, how they should be developed, and what they ultimately mean. As the episode demonstrates, abundance is easy; judgment remains the scarce creative resource.

Sources

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page