
Last Mile Agentic Image Editing
A Little Plains project.
Building a generative production workflow for product-dense imagery
Yesterday OpenAI released ChatGPT Image 2.0. Our teams using it have seen some marked improvement in text accuracy, but it is not a one-shot silver bullet for upscaling all the fine details for generative. This brings up a larger point on how do we complete production-grade generative work? How do you solve for the last mile?
A few months ago, @littleplainsxo was asked to assist a well-known 125 year old NYC-based home & furniture fixture with their Spring '26 editorial photo & video shoot. The catch is, we were asked to do it all 100% generatively.
We almost exclusively work with early-stage startups, but a mutual friend of ours was advising the business, knew how much we'd dived into generative node-based workflows for creative and connected us with the leadership. It was a great opportunity to see if generative-image workflows could produce rich, product-dense furniture scenes that still felt elevated, editorial, and true to the brand. So we dove in...
Firstly we hired a professional photographer who knew scene building, lighting, f-stops, everything technical. Then we paired them with @rameadows an AI design-engineer specialist, to work alongside our in-house all-things-generative expert @DanBatten and rethink how this could be done. Note: Rusty helped write this entire piece, thank you!
This project gave us a chance to see how these workflow could hold up inside a traditional ecommerce process with real creative and commercial scrutiny.
We were effectively bringing an emerging image pipeline “undercover” into a very established retail image system and learning what it takes for that workflow to perform at a high level.
This was not just one hero object, nor a clean little vignette. Full, layered scenes with rugs, furniture, lighting, tabletop, decor, and detail stacked together in a way that still felt editorial, believable, and beautifully composed.
That is what made the assignment both exciting and difficult. Little Plains was already working with generative tools, but not usually against exacting point-of-purchase expectations around product identity, materials, finishes, scale, lighting, and small details. Here, those details mattered a lot and we quickly learned where AI tools hit their limits.
What We Learned
This project helped us sharpen a workflow that combined generative tools, structured image direction, and traditional finishing to produce assets that could live across stills, motion, and social. Additionally, we needed to deliver all of this at a level of speed and flexibility that would have been difficult to match with a fully traditional shoot.
It was hard.
Models drifted. Small objects broke. Upholstery softened. Wood grain wandered. Perspective got slippery. Resolution ceilings showed up quickly. The final 10–20% routinely required manual intervention.
But that was the point of the project. We were not looking for a magic button. We were building a system.
The workflow was never deterministic. Even with good inputs, getting one usable hero often meant running 20–40 generations, adjusting image order, tightening spatial language, or rebuilding parts of the node tree until the scene finally held together.
The Challenge
Furniture imagery is unforgiving.
A great image has to do several jobs at once: sell the room, sell the mood, and sell the individual products inside it. That gets much harder when one scene contains rugs, sofas, chairs, lighting, side tables, tabletop objects, art, and decor all at once.
In a traditional shoot, that complexity is solved with location choices, physical styling, lighting design, camera decisions, retouching, and a lot of budget. In this workflow, those same demands still exist; they just move into a different stack. The problem becomes computational and editorial rather than purely physical. Every scene still needs composition, scale, contact, shadow logic, material truth, and product legibility. The client still expects the image to look resolved.
That was the real challenge here. We needed to make images that could survive review by a traditional ecommerce team that did not care how the image was made. If a chair leg bent, if a rug pattern drifted, if a lamp scale felt off, or if six products could not all read cleanly in a frame, the workflow had failed.
Dense scenes amplified every failure mode. Simple vignettes are relatively easy. A full room with six to ten products, multiple material categories, layered styling, and believable atmosphere is where the current tools start to show their seams.
Our Process
We approached the work as a repeatable production pipeline, not as one-off prompting.
1. Briefing and world-building
The client provided target personas and furniture collections. From that, we developed original interior environments from scratch for each persona: not yet product-heavy compositions, but plausible worlds with the right architecture, mood, camera logic, and stylistic constraints. These empty environments were reviewed and approved first. That step mattered because it separated environmental direction from product density and gave us a stable base to build on.
2. Product injection and hero generation
Once an environment was approved, we built structured prompts inside Flora to introduce the relevant furniture and decor into the scene. The goal was not just to place products, but to make them coexist inside one image with believable proportions, lighting response, and editorial composition. This is where the work got hard quickly. A hero image needed to satisfy atmosphere and accuracy at the same time.
Rather than progressively adding products one by one, we found the strongest hero images came from an “all at once” method. The approved empty room stayed fixed as Image 1, then rugs, furniture, lighting, tabletop, and decor assets were wired in together so the model had to solve global illumination, contact shadows, and spatial relationships across the full scene in a single pass.
In practice, that usually meant routing the empty room plus the ecommerce assets into a GPT 5.4 text node that generated structured JSON for Nano Banana Pro. We were very explicit about image order and spatial directives. A rug might be described as covering most of the floor plane of Image 1; a chair might need to sit three-quarters profile near the window; a lamp might need to read as lit but not glow unnaturally. We also started defining camera metadata directly in the prompt stack, including lens behavior and distortion constraints, because those decisions affected whether the image felt like a real interior photograph or a synthetic collage.
Nano Banana Pro was the main hero model because it handled the hardest part of the assignment best: multi-object grounding, room logic, and believable spatial physics. We typically rendered in 4K at a 3:2 landscape ratio and ran small batches rather than single shots, because lower-resolution generations broke down quickly when pushed further in finishing.
3. Iteration through feedback
The client reviewed hero options for overall composition, product rendering, alignment to the brief, and whether the image felt “right.” We then revised prompts, adjusted inputs, re-generated, and repeated. This loop was essential. The best images did not come from a single successful prompt; they came from rounds of correction and convergence.
This was also where model selection became tactical. Nano Banana Pro was the heavy-lift model for hero scenes, but faster support passes and text-sensitive fixes sometimes benefited from Nano Banana 2. The workflow was less about finding one perfect model than knowing which model was best suited to the current failure mode.
4. Precision finishing in Photoshop
Photoshop was the core infrastructure for the final leg of the project. Strong generations still needed product correction, edge repair, texture cleanup, compositional surgery, artifact removal, harmonization, selective color work, and upscaling. The last stretch was manual because current models still break in small but important ways. This is the most important operational truth from the project: at this complexity level, if the goal is a polished retail asset, there is currently no way around careful human finishing.
We worked non-destructively in layers and often combined the best elements from multiple generations into a single master file. If an object needed a full replacement, we could isolate the area, use Photoshop’s Generative Fill with a reference product shot, and manually blend the result back into the scene. In other cases, Flora or Nano Banana Pro could be used as a more natural-language correction layer to resize, re-texture, or re-ground a problem object. A final grain pass helped unify the polished composite and soften the “over-clean” digital look.
5. Alternate angles and video assets
Once a hero image was sufficiently locked, we used it as the basis for alternate-angle contact sheets. From those, we selected three to five additional views for review and refinement. Rather than relying on a single camera-move edit, we generated 2x2 contact sheets from the approved hero and explicitly restated styling props and room details so the model would not forget them in the alternate views. That let one approved world stretch across formats instead of stopping at a single still.
From those contact sheets, we extracted successful quadrants back into full stills instead of simply cropping and upscaling by hand. When needed, final stills were pushed through a separate upscale workflow to preserve detail while avoiding the mushiness that comes from enlarging weak generations.
6. Final assets
After the alt angles were landed, we went back to Photoshop to finalize the assets and export high resolution versions. This pass was a last check on color, harmony, and alignment across all project assets.
Final approved stills were then cropped into verticals for social and motion workflows and extended into short video clips through Flora using Veo 3.1 Ingredients. The video workflow used three clean 9:16 crops from the same still so the model could interpolate a slow, believable editorial pan rather than invent an entirely new shot. The best results came from long, gentle moves, repeated iterations, and removing the model’s scratch audio before final edit.
These videos along with final images for each hero and alternative angles were delivered to the client.
The Results
The images worked. More importantly, we found a workflow that worked.
We produced high-density furniture scenes that would have been difficult to achieve within the same budget and timeline through a traditional shoot. The value was not just lower cost. It was the combination of range, speed, and extensibility: once a world and hero composition were established, we could keep pushing the system into alternates, crops, and motion instead of starting over each time.
The process was not frictionless, and we do not want to pretend otherwise. We pushed many images as far as we could, but there was still a stopping point where additional perfection would have required disproportionate effort. That is part of the honest read on the current state of the tools. Even so, the client was happy with most of the final assets, and the project gave us a much sharper understanding of where the pipeline holds and where it needs reinforcement.
The most important result was not any single frame. It was proving that generative production can operate inside a traditional retail approval structure when it is handled rigorously. The unlock was not “AI replaces the shoot.” The unlock was “AI becomes a new production layer” for scenarios where building every asset physically would be too slow, too expensive, or too rigid.
Tools We Used
Throughout the project, we used a combination of tools from various vendors. In addition to the below tools we landed on, we explored and evaluated dozens of alternatives throughout the project.
Flora
@floraai was the main scene-generation and orchestration layer. We used it to build hero scenes with furniture in context, create alternate angles from approved compositions, test edits and tweaks, and generate motion from locked stills. GPT 5.4 text nodes inside Flora were especially useful for turning a stack of room and product references into structured JSON prompts rather than relying on loose natural-language prompting alone.
Photoshop
Photoshop was the precision layer. We used it for image correction, compositing adjustments, color harmonization, detail repair, spatial cleanup, upscale workflows, and final export preparation. If Flora gave us the world, Photoshop made the image hold together under scrutiny.
Gen AI Models
Nano Banana Pro was the primary still-image model for hero scenes and contact-sheet extraction. We used it when we needed the model to reason through room composition, product grounding, shadows, and multi-object physics all at once.
Nano Banana 2 was useful as a faster support model, especially when text rendering needed to hold more cleanly than the hero model could manage on its own. It was particularly helpful for problem areas like book spines, framed art, or other small text-bearing details.
Veo 3.1 was the motion extension model. We used Ingredients mode inside Flora to interpolate between cleaned vertical crops and turn a locked still world into a slow editorial pan for social video.
Note: ChatGPT's Image 2.0 is very helpful for last-mile text based accuracy. But, using node-based workflow tools is still highly recommended overall for production grade version-controlling.
Prompts & Techniques
One of the most useful parts of the workflow was moving away from loose natural-language prompting and toward structured prompting. We built a JSON-based prompt system inside Flora that let us define scenes more consistently and swap product sets more systematically. That gave us a more controllable way to manage room logic, camera behavior, styling constraints, and product insertion.
The exact schema changed over time, but the structure usually captured the same core ideas: what Image 1 represented, how the camera should behave, which products had to appear where, which props had to persist across angles, and what kind of image we were asking the model to return.
Example prompt structure:
Exact prompt structure:
{
"scene_goal": "Generate a photorealistic, high-end interior editorial image that feels like a real magazine photograph.",
"base_image": {
"image_1_role": "Approved empty room. Keep architecture, perspective, and lighting logic anchored to this image."
},
"camera": {
"lens": "40mm",
"distortion": "none",
"ratio": "3:2 landscape",
"framing": "eye-level interior photograph"
},
"style": {
"mood": "airy coastal minimal",
"lighting": "soft daylight with believable contact shadows",
"material_truth": "preserve wood grain, upholstery texture, and finish realism"
},
"placements": [
{
"image": "Image 2",
"object": "rug",
"instruction": "Cover most of the floor plane and anchor the main seating group."
},
{
"image": "Image 3",
"object": "sofa",
"instruction": "Place centered on rug, facing camera, scaled realistically to room."
},
{
"image": "Image 4",
"object": "accent chair",
"instruction": "Place near window in three-quarter view with believable floor contact."
}
],
"constraints": [
"Render all products in one pass rather than progressive compositing.",
"Keep shadows soft and consistent across all objects.",
"Do not invent extra furniture.",
"Return a resolved 4K hero image suitable for finishing in Photoshop."
]
}A few techniques mattered repeatedly:
- Approving environments before loading them with products.
- Treating hero images as systems to be stabilized, not lucky outputs to be accepted.
- Using an “all at once” hero method so the model solved lighting and physics globally.
- Using Nano Banana Pro for complex room logic and Nano Banana 2 when text fidelity mattered more.
- Using contact-sheet generations to explore alternate angles after composition lock.
- Extracting successful alternate panels into full stills instead of manually cropping weak source images.
- Separating exploration from finishing so the team could move quickly early and precisely late.
- Accepting that the final 10–20% often lives outside the model.
This is also where human judgment remained non-negotiable. The job was not primarily to have better “taste” than the model. The job was to notice failures: warped geometry, false joins, inconsistent materials, incorrect product identity, broken scale relationships, and lighting that looked plausible globally but wrong locally. Precision, not just aesthetics, was the differentiator.
What’s Next
The tool landscape is moving fast, and several pain points from this project are already improving. Expanded context windows, better spatial reasoning, sharper native resolution, and stronger object persistence should all make this workflow materially better over time.
But better models alone are not the full story. The next wave of improvement is also about inputs and control. We can make future runs stronger by normalizing product photography, gathering more angles per product, enforcing more consistent sizing and scale references, and tightening how assets are prepared before generation begins.
We also see a clear path toward hybrid workflows that combine generative models with lightweight 3D scaffolding. A low-resolution 3D environment could help lock camera, composition, layout, and scale before a model renders the scene. Product-aligned 3D or multi-angle references could also improve alternate-angle generation and reduce drift. In other words: the future is not pure prompting. It is better systems.
For Little Plains, that is the main takeaway. This project was less about making one set of images and more about stress-testing a new production method. It showed where generative tools are already useful, where they still break, and how a disciplined pipeline can bridge that gap today.
Thank you,
Emmett & Rusty @ Little Plains
