AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Training AI To Master Watercolour Art: Insights With TRL And OpenEnv on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A developer has recreated Surya Narreddi’s watercolour-generating AI model using open-source tools, releasing all artifacts publicly. The project tests reinforcement learning over aesthetic taste, with potential implications for AI art and interpretability.

An independent engineer has published a complete, open-source reproduction of Surya Narreddi’s watercolour-generating language model, utilizing TRL and OpenEnv frameworks on Hugging Face infrastructure. This project follows Narreddi’s viral August video, which showcased AI-produced watercolour paintings that gained over 1.5 million views, and aims to explore whether reinforcement learning can optimize models based on aesthetic preferences rather than traditional correctness.

The reproduction includes datasets, scripts, the reinforcement learning environment, and trained models, all openly available in a single Hugging Face collection. It implements a reward system combining four terms: a correctness gate, a length penalty, a style judge based on a vision-language model (Qwen3-VL-30B-A3B-Instruct), and a preference model (HPSv3). The training process involves 110 steps, 240 episodes per step, and uses Qwen/Qwen3.5-35B with LoRA adjustments, running end-to-end on Hugging Face infrastructure.

The core question of the project is whether reinforcement learning can be effectively applied to optimize AI models for aesthetic taste, moving beyond verifiable correctness to subjective quality. The project’s author emphasizes that this open release removes barriers for further research and experimentation, providing a baseline for future work in aesthetic AI training and interpretability.

At a glance
reportWhen: published March 2024
The developmentAn independent engineer has completed an open reproduction of Surya Narreddi’s viral watercolour AI model, deploying TRL and OpenEnv on Hugging Face infrastructure and releasing all related artifacts.
At a glance
reportWhen: published after the 23 August viral vid…
The developmentA fully open reproduction of Surya Narreddi’s viral watercolour-painting coding model — including the RL environment, reference dataset, training scripts and trained models — has been published, built with TRL and OpenEnv and running entirely on Hugging Face infrastructure.

Implications for AI Art and Aesthetic Optimization

This project demonstrates that reinforcement learning can be directed towards subjective aesthetic preferences, challenging the conventional focus on correctness or utility in AI training. It opens new avenues for AI-generated art, where models are tuned to produce visually appealing outputs aligned with human taste, rather than strictly functional or factual results. The open release of datasets, code, and models fosters transparency and collaborative development, potentially accelerating progress in AI art and interpretability. Moreover, the ability to inspect and edit code-based outputs offers a level of transparency absent in pixel-based image generators, making this approach relevant for artists, researchers, and developers interested in explainable AI.

Amazon

AI art generator software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Historical and Technical Background of AI Art Reproduction

The project situates itself within a lineage of generative AI art, from DeepDream (2015) to neural network portraits by Mario Klingemann and curated datasets by artists like Anna Ridler. Narreddi’s approach builds on earlier efforts to train models on curated datasets, but advances this tradition by employing reinforcement learning over aesthetic preferences rather than predefined correctness. His initial work focused on close-up flower images, but the current reproduction extends to full compositions, aiming to replicate the viral watercolour paintings created by Narreddi in August. The open-sourcing of all artifacts marks a significant step in democratizing access to such models and methodologies.

“This open reproduction provides a valuable resource for exploring how reinforcement learning can be used to optimize AI models for subjective aesthetic preferences.”

— Thorsten Meyer, AI researcher

Amazon

watercolour painting AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Quality and Effectiveness

It remains unclear how closely the reproduced outputs match the original viral paintings in terms of quality and stylistic fidelity. The project presents visual median outputs but lacks quantitative evaluation comparing different reward mixes. Furthermore, the full technical report from Narreddi, which is expected to provide deeper insights, has not yet been published. The effectiveness of reinforcement learning in truly capturing aesthetic preferences over multiple iterations is still under investigation, and the impact of the specific reward design remains to be validated through wider experimentation.

Amazon

artistic reinforcement learning models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Research Directions and Community Engagement

The next steps involve awaiting Narreddi’s full technical report for detailed evaluation and validation of the approach. The open artifacts enable other researchers and artists to replicate, modify, and extend the work, potentially testing different reward configurations or applying the methodology to other art styles. Further experiments may explore how well the model generalizes beyond the curated reference pool and whether reinforcement learning can be scaled to more complex or diverse artistic tasks. Community engagement and peer review will be crucial in assessing the robustness and artistic value of these models.

Amazon

visual style AI art software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main goal of this open reproduction?

The primary aim is to demonstrate that reinforcement learning can be used to optimize AI models for aesthetic preferences, moving beyond traditional correctness-based training, and to provide open tools for further research.

How does the reward system influence the model’s output?

The reward combines a correctness check, a preference score based on style comparisons, a length penalty, and a human-like preference model, guiding the model to produce visually appealing, stylistically consistent watercolour paintings.

Can this approach be applied to other art styles?

In principle, yes. The open framework allows for curating different reference datasets and adjusting reward components to target other styles or artistic goals, though further experimentation is needed.

What are the limitations of this project?

Uncertainties remain regarding how well the outputs match the original viral paintings in quality, and the effectiveness of RL over subjective taste has yet to be fully validated through rigorous evaluation.

How can I access the open artifacts?

All datasets, scripts, trained models, and environment configurations are published in a dedicated Hugging Face collection, available for download and experimentation.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Xiaomi: New CPU Matches Apple Cores Single Threaded, Much Faster Multithreaded

Xiaomi unveils a new CPU matching Apple cores in single-thread performance and significantly surpassing in multithreaded tasks, marking a major advancement.

Show HN: A Project Oberon System Version Running On RISC-V Instead Of RISC-5

A version of the Project Oberon operating system now runs on RISC-V hardware, replacing the original RISC-5 platform, as announced on Show HN.

‘GTA 6’ On PS5 Vs Xbox: Key Differences To Consider

Explore confirmed differences between GTA 6 on PlayStation 5 and Xbox, including graphics, performance, and exclusive features, as interest surges.

Spatial Focus Room: Make Distraction Impossible

A new app for Apple Vision Pro, Spatial Focus Room, aims to eliminate distractions by creating immersive environments for deep work, changing how focus is achieved.