I trained an AI clone of myself using ten photos and about ten minutes of processing. It can now generate photorealistic images of me anywhere on earth, in any wardrobe, for roughly the cost of a rounding error. Below are the exact prompts, six unedited results, what it actually cost, and the four places it still falls apart.

Summary

  • Input: 10 reference photos, varied angles and lighting. Training time: about 10 minutes.
  • Output: unlimited photorealistic images, roughly 2 to 4 credits each, which is cents per image.
  • Likeness holds well. Distinctive features like haircut, goatee, and build carry through reliably.
  • With a trained model you stop describing your face in prompts. You describe wardrobe, location, and light.
  • Real limits: one person per image, hands and text still glitch, and roughly half of generations get discarded.
  • The economics are absurd. A comparable travel shoot is thousands of dollars and weeks of logistics.
AI generated image of Cody Wise standing on a rain-slicked neon Tokyo street at night, a location the model placed him in without travel
Tokyo, at night, in the rain. I have never taken this photo. It took about forty seconds to make.

Table of Contents

  • What I actually did
  • The training set that worked
  • Six prompts and six results
  • What it costs
  • Where it breaks down
  • How to write prompts for a trained model
  • Should you do this?
  • FAQ

What Did I Actually Do?

I trained a likeness model, sometimes called a Soul or a character model, on photographs of myself. Once trained, it generates new images of me from a text description. No camera, no location, no photographer.

The full process took under half an hour, and most of that was choosing which photos to feed it. The training itself ran in about ten minutes while I worked on something else.

I did this for a boring practical reason. We publish several blog posts a day across two sites, and every post needs images. Stock photography looks like stock photography, and I was not going to fly to a different country every week to shoot header images. The interesting part is what happened once it worked, which I wrote about separately in how AI clones are changing social media and UGC.

What Training Set Actually Worked?

Ten photos. Variety matters far more than volume, and face visibility matters more than image quality.

What I includedWhy it helped
Front-facing headshotClean, well-lit baseline for facial structure
Sharp profile shotTeaches the side of the face and jaw line
Outdoor natural light portraitDifferent lighting conditions prevent a studio-only look
Indoor artificial lightSame reason, opposite direction
Several casual phone photosReal-world angles the posed shots do not cover

Two of my ten were weaker: shots where I was small in the frame. They did not break anything, but they contributed almost nothing. If I were doing it again I would use eight strong close-to-mid shots rather than pad the set.

The rule that matters: every photo should show the face clearly, and no two photos should be from the same session in the same light.

Six Prompts, Six Results

These are unedited first generations. No retouching, no cherry-picking from large batches. Each prompt follows the same skeleton: wardrobe, location, light, camera style.

1. Rocky Mountain cabin deck

Photorealistic editorial photograph, a man wearing a black tee and silver chain, working on a laptop on the wooden deck of a Rocky Mountain cabin, coffee mug and a golden retriever resting beside the chair, pine forest and snow-capped peaks behind, soft morning light and mist, candid documentary style, shallow depth of field, warm natural color, no readable text on screen
AI generated image of Cody Wise working on a laptop from a Rocky Mountain cabin deck, produced from a trained likeness model rather than a photoshoot
The dog was not in the training data. I asked for one and got one.

2. Tokyo side street at night

Photorealistic editorial photograph, a man wearing a black bomber jacket and black cap, standing on a neon-lit Tokyo side street at night in light rain, reflections on wet asphalt, vending machines and izakaya signs glowing behind him, cinematic street photography, shallow depth of field, moody color grade

Result is the image at the top of this article. Note what the model handled without being told: the wet-asphalt reflections, the colour temperature mismatch between neon and ambient, and the depth falloff down the street. That is the part that would have cost real money to shoot.

3. Banff chairlift

Photorealistic editorial photograph, a man wearing black ski gear and goggles pushed up on his helmet, sitting on a chairlift with snow-covered Banff mountain terrain below, bright winter sun, breath visible in cold air, candid action documentary style, shallow depth of field, crisp cool color
AI generated image of Cody Wise on a chairlift above snow-covered Banff terrain, showing the model handling gear and cold-weather lighting
Full gear change plus harsh snow-reflected light. Both are hard, both held.

4. Boat off a Croatian island

Photorealistic editorial photograph, a man with a very short buzz cut and a high skin fade on the sides and a well-groomed short dark beard with clean sharp edges, wearing a white linen shirt open over swim shorts and sunglasses, sitting on the bow of a small wooden boat anchored in turquoise water off a Croatian island, sun-bleached limestone cliffs behind, bright midday Mediterranean light, candid travel documentary style, shallow depth of field
AI generated image of Cody Wise on a boat in turquoise Adriatic water, demonstrating the model in bright natural midday light
Harsh overhead midday sun is the least forgiving light there is. This is the one I expected to fail.

5. Desert highway

Photorealistic editorial photograph, a man in a black New York Yankees fitted cap with a well-groomed short dark beard with clean sharp edges, wearing a black tee and sunglasses, leaning against a vintage convertible parked on a desert highway pullout at golden hour, red rock formations and open road stretching behind, warm dusty light, cinematic travel documentary style, shallow depth of field
AI generated image of Cody Wise leaning against a convertible on a desert highway at golden hour, a scene built entirely from a text prompt
Sunglasses partially occlude the face and the likeness still reads.

6. Rome at dawn

Photorealistic editorial photograph, a man in a black New York Yankees fitted cap with a well-groomed short dark beard with clean sharp edges, wearing a charcoal wool overcoat and scarf, walking through a narrow cobblestone street in Rome at dawn, espresso bar awning and parked scooters lining the street, soft golden early light, candid street photography, shallow depth of field, subtle film grain
AI generated image of Cody Wise walking a cobblestone Rome street at dawn, showing the model in a full winter wardrobe change
A wardrobe I do not own, in a city I was not in, at an hour I would not have been awake.

What Does This Actually Cost?

Generation runs roughly 2 to 4 credits per image depending on resolution, which lands at cents per image on a normal subscription. Training was a one-time cost. The whole set above cost less than a coffee.

Traditional shootTrained likeness model
SetupPhotographer, location, travel, scheduling10 photos, one upload
Lead timeDays to weeksAbout 10 minutes
Cost per usable imageTens to hundreds of dollarsCents
Location rangeWherever you can physically goAnywhere describable
AuthenticityRealNot real, and must be disclosed

That last row is the one that matters, and it is why this is a production tool rather than a replacement for showing up.

Where Does It Break Down?

Four real limits, none of which the demos usually mention.

  1. One person per image. A trained likeness handles a single subject. Anything with you and another specific person needs a different technique entirely.
  2. Hands and text still glitch. Fingers occasionally go wrong and any small text in frame comes out as convincing nonsense. Always include an instruction to avoid readable text, and always check hands before publishing.
  3. Roughly half of generations get discarded. Not because they are broken, but because something is slightly off: a strange expression, an odd proportion, a detail that reads wrong. Budget for curation.
  4. Fine personal detail drifts. Tattoo placement is approximate. Specific expressions tied to real moments do not reproduce. It gets the person, not the particular.

And the hard line: this cannot be used for anything that implies proof. Not testimonials, not results claims, not evidence you were somewhere. Regulators treat synthetic testimonials as deceptive whether or not you label them.

How Do You Write Prompts for a Trained Model?

Stop describing your face. That is the single biggest change, and it is counterintuitive if you are used to normal image generation.

Before training, a prompt needs the full physical description: age, hair, skin tone, facial hair, build. After training, all of that is baked in. Repeating it actively fights the model and produces worse likeness.

Prompt elementInclude?
Facial features, hair, skin toneNo. The model supplies these
WardrobeYes. Be specific
Location and background detailYes. This is most of the work
Lighting and time of dayYes. Biggest driver of realism
Camera and framing styleYes. Controls whether it reads editorial or stock
Negative constraintsYes. “no readable text on screen” earns its place

The skeleton I use: photorealistic editorial photograph, a man wearing [wardrobe], [action] [location], [lighting], candid documentary style, shallow depth of field, warm natural color, no readable text on screen. Swap the bracketed parts and everything else stays constant, which is what keeps a brand looking consistent across dozens of posts.

Should You Actually Do This?

Yes if you publish frequently and your face is part of the brand. The maths is not close: consistent on-brand imagery at near-zero marginal cost solves a real bottleneck for anyone shipping content daily.

No if you were hoping it replaces being visible. It generates pictures of you. It does not generate credibility, and it cannot do the one thing your audience is actually evaluating, which is whether you know what you are talking about.

The honest framing: this is a production tool that removes a logistics problem. It is the same category as a good camera or a design system, not a shortcut around having something worth saying. If you want help building the content and brand system underneath, our branding packages and website growth packages cover the part that no model generates for you.

Frequently Asked Questions

How many photos do you need to train an AI clone?

Between 5 and 20. Ten is comfortable. Variety and face visibility matter more than quantity: different angles, lighting, and distances, with the face clearly visible. Shots where you are small in frame contribute very little.

How long does training take?

Roughly ten to fifteen minutes of processing. Selecting and cleaning the reference photos takes longer than the training.

How much does it cost to generate AI brand photos?

Around 2 to 4 credits per image depending on resolution, which is cents per image on a normal subscription. Against a photoshoot with a photographer, location, and travel, the per-image difference is orders of magnitude.

Do AI generated photos actually look like you?

With a trained model, yes. People who know you recognise you immediately. Distinctive features like haircut, facial hair, and build carry reliably. Fine detail such as exact tattoo placement is where accuracy drops.

What are the limitations?

One person per generation, occasional hand and text artefacts, about half of outputs discarded, and drift on fine personal detail. It also cannot be used for testimonials or anything implying proof.

Is it legal to use AI photos of yourself in marketing?

Using your own likeness in your own marketing is generally straightforward since you control your image rights. Platform disclosure obligations still apply. Cloning anyone else, including employees or hired creators, requires explicit written licensing.

The Bottom Line

Ten photos and ten minutes bought unlimited on-brand photography of a person who was sitting at a desk in Calgary the entire time. The production constraint that shaped content strategy for two decades is simply gone.

What has not changed is that a picture of you standing in Tokyo does not make you more worth listening to. The tool removes an excuse, not a requirement. Use it to publish more consistently, not to pretend to a life you are not living.

Every image in this article was generated by AI from a trained likeness model. If you want a content system that uses this properly, tell us where you are at through our project intake form.