AI Image & Video

Lesson 2 of 16

The main tools

The AI visual world has a handful of major tools, each with strengths. Knowing the main tools — for images and for video — helps you pick the right one and get started without being overwhelmed by the fast-moving landscape. This lesson covers the main tools. Let's map the toolkit. Let's get oriented.

The main image tools

The AI visual landscape has several MAJOR TOOLS. They change + improve constantly (new ones appear,
existing ones update monthly), so focus on the CATEGORIES + leading players + the CONCEPTS (which
transfer) rather than memorising a fixed list. Grouped by what they do:

AI IMAGE GENERATION TOOLS (the big ones):
- MIDJOURNEY — renowned for STUNNING, artistic, high-quality images. A favourite for beautiful,
  stylised results. (Historically via Discord; now also a web app.) Great aesthetic quality; a slight
  learning curve. Paid.
- DALL·E (OpenAI) — integrated into ChatGPT — very easy to use via natural conversation; good at
  following prompts + handling text-in-images. Accessible + beginner-friendly.
- STABLE DIFFUSION — an OPEN-SOURCE model you can run yourself (free, locally) or via services. Highly
  CUSTOMISABLE + controllable (custom models, extensions, fine-tuning via tools like Automatic1111/
  ComfyUI). The power-user/technical choice; endless flexibility.
- ADOBE FIREFLY — Adobe's generative AI, built into Photoshop + trained on licensed content
  (commercially safer). Great for editing (Generative Fill) + integration with pro design tools.
- GOOGLE IMAGEN (Gemini), LEONARDO.AI, IDEOGRAM (great at text in images), + others — a growing field.
- CANVA + built-in AI — many design tools (Canva, etc.) now embed AI image generation — convenient for
  non-technical users making designs.
CHOOSING: MIDJOURNEY for artistic beauty, DALL·E/ChatGPT for ease, STABLE DIFFUSION for control/free,
FIREFLY for editing + commercial safety, Canva for all-in-one design. Try a few; find your favourites.

The AI visual landscape has several major tools — they change + improve constantly, so focus on the categories + leading players + the concepts (which transfer) rather than memorising a fixed list. AI image generation tools (the big ones): Midjourney (renowned for stunning, artistic, high-quality images — a favourite for beautiful, stylised results — historically via Discord, now also a web app — great aesthetic quality; a slight learning curve; paid); DALL·E (OpenAI) (integrated into ChatGPT — very easy to use via natural conversation; good at following prompts + handling text-in-images — accessible + beginner-friendly); Stable Diffusion (an open-source model you can run yourself (free, locally) or via services — highly customisable + controllable (custom models, extensions, fine-tuning via Automatic1111/ComfyUI) — the power-user/technical choice; endless flexibility); Adobe Firefly (Adobe's generative AI, built into Photoshop + trained on licensed content (commercially safer) — great for editing (Generative Fill) + integration with pro design tools); Google Imagen (Gemini), Leonardo.ai, Ideogram (great at text in images) (a growing field); and Canva + built-in AI (many design tools now embed AI image generation — convenient for non-technical users) — choosing: Midjourney for artistic beauty, DALL·E/ChatGPT for ease, Stable Diffusion for control/free, Firefly for editing + commercial safety, Canva for all-in-one design — try a few; find your favourites.

The main video tools and doing it well

AI VIDEO TOOLS (a fast-evolving frontier):
- TEXT/IMAGE-TO-VIDEO GENERATORS:
  - RUNWAY (Gen-3/Gen-4) — a leading creative AI video tool (text-to-video, image-to-video, + editing
    features). Popular with creators.
  - OPENAI SORA — high-quality, longer, coherent text-to-video (impressive realism).
  - GOOGLE VEO, KLING, PIKA, LUMA (Dream Machine), MINIMAX/Hailuo — strong text/image-to-video
    generators, each improving fast. Many offer free trials.
  - Use for: short clips, animated scenes, b-roll, social content, concept videos.
- AI AVATAR / TALKING-HEAD VIDEO:
  - HEYGEN + SYNTHESIA — create realistic AVATAR videos that speak your script (for training,
    marketing, explainers) with AI voices — no filming. D-ID too.
  - Great for talking-presenter content without a camera/actor.
- AI VOICE / AUDIO:
  - ELEVENLABS — leading realistic AI VOICE generation (voiceovers, narration, cloning). Play.ht,
    Murf too.
  - Add voiceovers to your videos without recording.
- AI-ENHANCED VIDEO EDITING:
  - CAPCUT, DESCRIPT, ADOBE (Premiere/After Effects with AI features) — edit + enhance video with AI
    (auto-captions, editing by text, effects, background removal). For assembling + polishing.

DOING IT WELL:
- START WITH ONE OR TWO TOOLS — don't try everything. Pick a main IMAGE tool (e.g. Midjourney or DALL·E/
  ChatGPT or a free one) + a main VIDEO tool (e.g. Runway or a free trial) + learn them well. Depth
  over breadth.
- MATCH THE TOOL TO THE JOB — artistic images -> Midjourney; easy/conversational -> DALL·E; control/free
  -> Stable Diffusion; editing/commercial -> Firefly; video clips -> Runway/Sora/Kling; talking videos
  -> HeyGen/Synthesia; voice -> ElevenLabs.
- MIND COST + FREE TIERS — many offer free trials/credits; try before paying. Paid tools give better
  quality/more use.
- FOCUS ON TRANSFERABLE SKILLS — prompting, composition, iterating, editing (this course) apply across
  ALL tools. Learn the craft, not just one app's buttons — so you adapt as tools change.
- EXPECT CHANGE — the tools evolve monthly. Stay curious + adaptable; the leader today may differ
  tomorrow. Concepts endure; specific tools shift.

THE PRINCIPLE: know the MAIN TOOLS by CATEGORY — IMAGE generation (MIDJOURNEY artistic, DALL·E/ChatGPT
easy, STABLE DIFFUSION free/controllable, FIREFLY editing/commercial, Canva all-in-one), VIDEO
generation (RUNWAY, SORA, VEO, KLING, PIKA), AI AVATARS (HEYGEN/SYNTHESIA), + AI VOICE (ELEVENLABS).
Pick ONE OR TWO to master, MATCH the tool to the job, use free tiers, + focus on TRANSFERABLE skills
(the tools change fast; the craft endures). Get oriented, then dive deep into a couple.

AI video tools (a fast-evolving frontier): text/image-to-video generators (Runway (Gen-3/Gen-4) — a leading creative AI video tool; OpenAI Sora — high-quality, longer, coherent text-to-video; Google Veo, Kling, Pika, Luma (Dream Machine), MiniMax/Hailuo — strong generators, each improving fast, many with free trials — use for short clips, animated scenes, b-roll, social content); AI avatar / talking-head video (HeyGen + Synthesia — realistic avatar videos that speak your script, with AI voices, no filming; D-ID too — great for talking-presenter content without a camera); AI voice / audio (ElevenLabs — leading realistic AI voice generation (voiceovers, narration, cloning); Play.ht, Murf too — add voiceovers without recording); and AI-enhanced video editing (CapCut, Descript, Adobe (Premiere/After Effects with AI) — edit + enhance with AI: auto-captions, editing by text, effects, background removal — for assembling + polishing). Doing it well: start with one or two tools (don't try everything — pick a main image tool + a main video tool + learn them well — depth over breadth); match the tool to the job (artistic → Midjourney; easy → DALL·E; control/free → Stable Diffusion; editing/commercial → Firefly; video → Runway/Sora/Kling; talking videos → HeyGen/Synthesia; voice → ElevenLabs); mind cost + free tiers (try before paying; paid tools give better quality/more use); focus on transferable skills (prompting, composition, iterating, editing apply across all tools — learn the craft, not just one app's buttons); and expect change (the tools evolve monthly — stay curious + adaptable; concepts endure, specific tools shift). The principle: know the main tools by category — image generation (Midjourney artistic, DALL·E/ChatGPT easy, Stable Diffusion free/controllable, Firefly editing/commercial, Canva all-in-one), video generation (Runway, Sora, Veo, Kling, Pika), AI avatars (HeyGen/Synthesia), + AI voice (ElevenLabs); pick one or two to master, match the tool to the job, use free tiers, + focus on transferable skills (the tools change fast; the craft endures).

The mistake beginners make

The first mistake is tool-hopping (trying everything) — spreading thin across dozens of tools, mastering none; pick one or two + go deep. The second mistake is using the wrong tool for the job — forcing an easy tool for artistic work (or vice versa); match the tool to the job. The third mistake is learning buttons, not craft — memorising one app's interface instead of transferable skills (prompting/composition/editing) that apply everywhere; learn the craft. And paying before trying — buying subscriptions without using free tiers first; try free before paying. And panicking at the fast change — feeling overwhelmed by monthly new tools; focus on concepts (they endure) + stay adaptable. Master one or two, match tool to job, learn the craft, use free tiers, and focus on enduring concepts.

Your turn


Your turn

  1. Know the main IMAGE tools: Midjourney (stunning artistic quality), DALL-E/ChatGPT (easiest, conversational), Stable Diffusion (open-source, free, highly customisable), Adobe Firefly (editing + commercial-safe, in Photoshop), plus Leonardo/Ideogram/Canva AI.
  2. Know the main VIDEO tools: text/image-to-video generators (Runway, OpenAI Sora, Google Veo, Kling, Pika, Luma), AI avatar/talking-head (HeyGen, Synthesia), AI voice (ElevenLabs), and AI-enhanced editing (CapCut, Descript, Adobe).
  3. Pick one or two to master: choose a main image tool + a main video tool (using free tiers/trials) and learn them well - depth over breadth, rather than tool-hopping.
  4. Match the tool to the job: artistic images -> Midjourney; easy/conversational -> DALL-E; control/free -> Stable Diffusion; editing/commercial -> Firefly; video clips -> Runway/Sora/Kling; talking videos -> HeyGen/Synthesia; voiceover -> ElevenLabs.
  5. Focus on transferable skills: prioritise prompting, composition, iterating, and editing (this course) - which apply across ALL tools - over memorising one app's buttons, since the specific tools change fast but the craft endures.

Key points

  • Know the MAIN TOOLS by CATEGORY (they change monthly — focus on categories + concepts, not a fixed list). IMAGE generation: MIDJOURNEY (stunning artistic quality — paid), DALL·E/ChatGPT (easiest, conversational, good text-in-image), STABLE DIFFUSION (open-source, FREE/local, highly customisable — power users), ADOBE FIREFLY (editing/Generative Fill + commercially safer — in Photoshop), Leonardo/Ideogram/Canva AI.
  • VIDEO generation: RUNWAY (leading creative video), OpenAI SORA (high-quality/longer), Google VEO, KLING, PIKA, LUMA — text/image-to-video, improving fast, many free trials. AI AVATARS (talking-head): HEYGEN + SYNTHESIA (avatar speaks your script — no filming). AI VOICE: ELEVENLABS (realistic voiceover/narration). AI EDITING: CapCut/Descript/Adobe (auto-captions, edit-by-text, effects).
  • CHOOSE by job: MIDJOURNEY = artistic beauty, DALL·E = ease, STABLE DIFFUSION = control/free, FIREFLY = editing/commercial, RUNWAY/SORA/KLING = video, HEYGEN/SYNTHESIA = talking videos, ELEVENLABS = voice.
  • Do it well: START with one or two tools (depth over breadth), MATCH the tool to the job, mind COST + use FREE TIERS/trials before paying, focus on TRANSFERABLE skills (prompting/composition/iterating/editing apply across ALL tools — learn the craft, not one app's buttons), and EXPECT CHANGE (tools evolve monthly; concepts endure).
  • The mistakes: tool-hopping (master one or two), using the wrong tool for the job, learning buttons not craft, paying before trying (use free tiers), and panicking at the fast change (focus on enduring concepts + stay adaptable).

Q&A · 0

Enrol to ask questions and join the discussion.

No questions yet — be the first to ask.