"I built a pipeline that takes a plain-English feature description, uses a Gen AI model to generate structured test cases from it, then I convert the highest-value cases into automated Selenium tests that run on every code push through a CI pipeline in GitHub Actions."
Manually writing every test case for every feature doesn't scale. But blindly trusting AI-generated tests is worse — you still need a human to judge which cases actually matter. This project shows both halves: using AI to speed up the thinking part of testing, and still applying judgment + automation skill to the execution part.
Feature description (plain English)
|
v
[1] generate_test_cases.py --> calls Gen AI API (OpenAI/Claude)
|
v
test_cases.csv / .json (structured: ID, Type, Steps, Expected Result)
|
v
[2] I manually review & pick the highest-priority cases
|
v
[3] Selenium + TestNG automated test scripts (Java)
|
v
[4] GitHub Actions CI pipeline
- runs automatically on every git push / pull request
- fails the build if a test fails
|
v
[5] Test report (pass/fail results, logs)
ai-qa-project/
├── ai-test-generator/
│ ├── generate_test_cases.py # calls the Gen AI API to generate test cases
│ ├── feature.txt # example input: feature description
│ ├── requirements.txt
│ └── sample_output/
│ └── test_cases.csv # example AI-generated output
├── selenium-tests/
│ ├── pom.xml # Maven project (Java + Selenium + TestNG)
│ ├── testng.xml
│ └── src/test/java/tests/
│ └── LoginTest.java # automated test converted from an AI-generated case
└── .github/workflows/
└── ci.yml # CI pipeline: runs tests on every push
generate_test_cases.py sends a feature description to a Gen AI API and asks it
to return structured test cases (positive, negative, and boundary cases) as
JSON. If no API key is configured, it falls back to a demo mode using a
pre-written example response — this is what's in sample_output/test_cases.csv
so you can see the output shape without needing an API key.
Example input (feature.txt):
"User login with email and password. Account locks after 3 failed attempts."
Example AI output (simplified):
| Test ID | Type | Steps | Expected Result |
|---|---|---|---|
| TC01 | Positive | Enter valid email + password, click login | User is redirected to dashboard |
| TC02 | Negative | Enter valid email, wrong password | Error message shown, no login |
| TC03 | Boundary | Enter wrong password 3 times | Account is locked, further attempts blocked |
I don't automate every AI-generated case. I review the list and pick the ones that test real risk — this is the "judgment" part interviewers care about, not just "I used AI."
LoginTest.java is a Selenium + TestNG script that automates TC01–TC03 against
a public demo login page (saucedemo.com, a standard practice site for QA
automation). It's written in Java to match my existing Java background.
.github/workflows/ci.yml runs the Selenium/TestNG suite automatically:
- On every
git pushand pull request - Using Maven to build and run tests
- Fails the pipeline (red X on GitHub) if any test fails, just like a real QA gate in a dev team's workflow
TestNG generates an HTML report (test-output/) showing pass/fail per test,
which is uploaded as a build artifact in the CI run.
AI test case generator:
cd ai-test-generator
pip install -r requirements.txt
export OPENAI_API_KEY=your_key_here # optional — omit to use demo mode
python generate_test_cases.pySelenium tests:
cd selenium-tests
mvn testCI pipeline: runs automatically on GitHub once you push this repo — no manual step needed. Check the "Actions" tab on GitHub.
Q: Why not just automate 100% of the AI-generated cases? A: Because not every generated case is worth the maintenance cost of an automated script. I prioritize cases with the highest risk or most user impact, and leave the AI's low-value or duplicate suggestions as manual/exploratory notes instead.
Q: What happens if the AI generates a wrong or nonsensical test case? A: That's exactly why step 2 (human review) exists. I treat the AI's output as a first draft, not ground truth — same way I'd treat a junior tester's first pass.
Q: What would you improve if you kept building this? A: I'd connect it to actual application logs so the AI could analyze failures and suggest new edge cases based on real production bugs, closing the loop between testing and monitoring. I'd also add Cypress support for pure UI/JS apps and Postman/Newman for API-level test generation.
Q: Why Selenium instead of Cypress? A: I used Selenium here because I'm most comfortable in Java, but the architecture (AI generates cases → human filters → automate → CI/CD) is tool- agnostic — the same pipeline works with Cypress, Playwright, or Postman/Newman for API testing.