Does editing text remove an AI flag?
Depends entirely on what you edit. Word-level changes move a score very little. Structural changes move it a lot, and most people spend their time on the first kind.
The intuitive approach is to go through and change words. It is the least effective thing available, and it is what almost every tool marketed for this purpose does.
Why word-level editing does so little
Detectors read two properties, and only one of them is about words.
Word choice predictability does respond to vocabulary changes, but weakly, because replacing a common word with a less common one in a sentence that keeps the same shape leaves the surrounding structure intact.
Sentence rhythm does not respond at all. Swapping words inside a sentence leaves the sentence the same length, so the variation across the document is unchanged. That variation is the strongest single signal most detectors use.
Human academic writing in our corpus varies sentence length with a standard deviation of 8.7 words, generated text with 6.4. An afternoon of synonym substitution moves that number by essentially nothing.

The editing that moves it the wrong way
Worth knowing, because this is the trap.
Polishing. A general tidy-up smooths prose, and smoothing removes variation. Text that goes through a careful edit comes out flatter than it went in, so the score goes up.
This is why a fifth draft scores higher than a first, and why people who edit hard in response to a flag sometimes make it worse. See why good writing gets flagged more.
Paraphrasing tools. Same problem with an extra layer: a tool applies one uniform transformation across the whole document, replacing whatever irregularity was there with its own consistency. See does QuillBot get flagged.
The editing that works
Three changes, all structural.
Restructure sentences, do not reword them. Split a long sentence into two. Merge two short ones. Put a four-word sentence next to a thirty-word one. This changes the property that carries most of the signal, and it takes minutes rather than hours.
Delete connective openings. "Moreover", "Furthermore", "Additionally", "Overall". These open 2.4% of human sentences and 5.8% of generated ones. Most can go with nothing lost.

Add back the specific. Replace "various approaches have been proposed" with two named approaches. Replace "significant improvement" with the number. Generalisation is the habit that reads as generated, and it is also usually a weaker sentence.
Break the paragraph template. Generated paragraphs run to a shape: topic sentence, support, restatement. Let one paragraph be two sentences. Let another run long.
Full list in what makes text read as machine written.
How to tell whether it worked
Measure it rather than guessing, per paragraph, before and after. In most documents two or three paragraphs carry the whole signal, so the work is smaller than it looks once you can see which ones.
Ours is free and unlimited with no account, takes .pdf, .docx and .tex, has no word cap so you can re-run as many times as you like, scores every paragraph separately, and never stores your text. Wrong-flag rate 0.4%, published, against Turnitin's 4%.