AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Breaking Down AI Watercolour Painting With TRL And OpenEnv on ThorstenMeyerAI.com

TL;DR

A developer has created an open, end-to-end reproduction of Surya Narreddi’s viral watercolour AI model, utilizing TRL and OpenEnv. All artifacts, including datasets, training scripts, and models, are now publicly available, enabling further research into aesthetic reinforcement learning.

An independent engineer has published a comprehensive, open-source reproduction of Surya Narreddi’s viral watercolour painting language model, making all datasets, training scripts, and trained models freely accessible on Hugging Face. This development allows the community to explore reinforcement learning focused on aesthetic preferences rather than verifiable answers, marking a significant step in AI art research.

The project replicates Narreddi’s model that gained over 1.5 million views in August, which used a language model to generate JavaScript code for watercolour-style paintings via p5.js. For more details, see the original analysis on training a coding model to paint watercolours. The reproduction employs TRL and OpenEnv frameworks to train a Qwen-based model with reinforcement learning against a composite reward function. This reward combines code correctness, style adherence, and human preference scores derived from a model trained on human choices. This approach is discussed in detail in the analysis of recent AI investments. All training environments, datasets, and models are openly shared on Hugging Face, enabling others to verify, modify, or extend the work. Learn more about open-source AI projects at this coverage of AI investments. Unlike typical reinforcement learning that optimizes for clear-cut answers, this project explores ‘RL over taste,’ where aesthetic judgment guides the training process, reflecting a shift towards more subjective, human-like preferences in AI art. The original project was created by Surya Narreddi, who described the process as adding natural drawing tools to p5.js, with the model producing code that can be read and edited, thus offering transparency in artistic decision-making. The reproduction’s release aims to foster community engagement and further investigation into aesthetic reinforcement learning, a relatively underexplored area compared to traditional, verifiable reward-based AI training.

At a glance
breakingWhen: announced March 2024
The developmentAn independent engineer has publicly released a complete reproduction of Surya Narreddi’s viral watercolour AI model, including datasets, training code, and models, on Hugging Face.
At a glance
reportWhen: published after the 23 August viral vid…
The developmentA fully open reproduction of Surya Narreddi’s viral watercolour-painting coding model — including the RL environment, reference dataset, training scripts and trained models — has been published, built with TRL and OpenEnv and running entirely on Hugging Face infrastructure.

Implications for AI Art and Reinforcement Learning

This open reproduction represents a major advance in AI art research by providing transparent, replicable artifacts that enable community experimentation with aesthetic reinforcement learning. It challenges the dominance of pixel-based image models by focusing on code-based, editable outputs, which allow for inspection and modification of each artistic decision. The project’s emphasis on aesthetic preferences rather than objective correctness could influence future AI creative tools, making them more aligned with human taste and subjective judgment. Additionally, the open release lowers barriers for researchers and artists to explore how reinforcement learning can optimize for beauty, style, or artistic coherence, potentially leading to new forms of AI-generated art that are more diverse and expressive. Overall, this effort highlights the importance of open science in AI art, fostering innovation and transparency in a field often characterized by proprietary models and opaque training processes.

Amazon

watercolour AI art software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Technical Foundations of the Reproduction

The original project by Surya Narreddi, which went viral in August, involved training a language model to generate JavaScript code that produces watercolour-style paintings through the p5.js library. This approach offers interpretability, as each brushstroke decision is explicit in the code, contrasting with pixel-based image generators that produce opaque outputs. Narreddi’s work is rooted in the tradition of AI art pioneers like DeepDream, GAN portraits, and neural-network-based artworks, emphasizing the medium’s artistic potential. The initial model was trained using a reward structure combining code correctness, stylistic fidelity, and human preferences, with the latter derived from a model trained on curated human choices. The open reproduction expands on this by providing all datasets, training scripts, and models, enabling others to verify and build upon the original research. The use of TRL and OpenEnv frameworks facilitates reinforcement learning against subjective aesthetic goals, a relatively novel direction in AI training, which historically focused on objective measures such as accuracy or code passing tests.

“This reproduction aims to make the entire process transparent and accessible, allowing the community to explore how reinforcement learning can be applied to aesthetic judgment.”

— Thorsten Meyer, project author

Amazon

digital painting tablet for AI art

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Performance and Fidelity

While all datasets, scripts, and models are now publicly available, it is still unclear how closely the reproduction’s outputs match the original viral paintings in quality and style. The project compares different reward mixes visually but does not provide a definitive quantitative assessment of which approach yields the best results. Additionally, the full technical report from Narreddi, which promises to clarify technical details and evaluation metrics, has not yet been published. It remains uncertain how the community will validate or extend this work, and whether the aesthetic preferences captured by the reward functions generalize beyond the curated reference pool.

Amazon

p5.js coding kit for artists

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Community Engagement and Research Validation

The immediate next step is for the community to test and evaluate the reproduction by experimenting with the shared datasets and models. Researchers may compare the outputs against original artworks or develop new reward functions to explore different aesthetic criteria. The release of all artifacts invites collaborative validation and improvement, potentially leading to more refined models that better capture human artistic preferences. Additionally, the upcoming technical report from Narreddi will likely provide further insights into the training process, evaluation methods, and plans for future work. As the field advances, further studies could investigate how reinforcement learning over taste influences creativity, diversity, and artistic expression in AI-generated art.

Amazon

AI art reinforcement learning books

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is included in the open reproduction release?

The release includes datasets, training scripts, the trained models, and the reinforcement learning environment, all hosted on Hugging Face for community access and experimentation.

How does reinforcement learning over aesthetic taste differ from traditional methods?

Traditional reinforcement learning relies on verifiable, objective rewards like correctness or passing tests. This project uses a reward based on human preferences and style, which are subjective and not strictly verifiable, representing a shift toward more human-like aesthetic optimization.

Can I reproduce the original viral watercolour paintings with this model?

The reproduction aims to replicate the process and outputs, but it is not yet confirmed how closely the generated paintings match the original viral artworks in style or quality. Community testing will help clarify this.

Will the original technical report from Narreddi be published?

Yes, Narreddi has stated that a full technical report is forthcoming, which is expected to provide detailed insights into the training process, evaluation metrics, and future directions.

What are the broader implications of this work?

This project demonstrates the potential for AI to learn aesthetic preferences through reinforcement learning, opening new avenues for creative AI applications that prioritize style, beauty, and subjective taste over objective correctness.

Primary source: Hugging Face · via ThorstenMeyerAI.com

You May Also Like

Ukraine’s Digital Warfare Tactics Enhanced By Artificial Intelligence

Ukraine integrates artificial intelligence into its cyber and logistical operations to disrupt Russian supply networks, marking a new phase in digital warfare.

The August 1 Cutoff: Making AI Benchmarks A Hidden Security Asset

The US government sets an August 1 deadline to establish classified AI benchmarks and voluntary pre-release review processes, impacting AI security and industry practices.

Origin Lab raises $8M to help video game companies sell data to world-model builders

Startup Origin Lab secures $8 million in seed funding to connect video game assets with AI research labs for training world models, benefiting both industries.

The Case For MUDs In Modern Times (2018)

Exploring why Multi-User Dungeons remain relevant in 2018, examining their benefits, challenges, and potential future in modern gaming.