Comments (4)
Same here.. Can you plz let us know if this is possible to be done from the
api?
cheers and congrats for the excelence work!
D.
Original comment by [email protected]
on 23 Nov 2011 at 11:18
from boilerpipe.
I searched litle bit more and i found the solution:
1)if your input is a string
private final BoilerpipeExtractor extractor =
CommonExtractors.DEFAULT_EXTRACTOR;
private final HTMLHighlighter hh = HTMLHighlighter.newExtractingInstance();
InputSource is = new InputSource(new StringReader(detailPageSourceCode));
final TextDocument doc = new BoilerpipeSAXInput(is).getTextDocument();
extractor.process(doc);
StringBuilder bf = new StringBuilder();
bf.append("<meta http-equiv=\"Content-Type\" content=\"text-html;
charset=utf-8\" />");
bf.append(hh.process(doc, detailPageSourceCode));
2)if your input is a URL(taken from HTmlHighlighterDemo.java)
URL url = new URL(
"http://research.microsoft.com/en-us/um/people/ryenw/hcir2010/challenge.html"
// "http://boilerpipe-web.appspot.com/"
);
// choose from a set of useful BoilerpipeExtractors...
final BoilerpipeExtractor extractor = CommonExtractors.ARTICLE_EXTRACTOR;
// choose the operation mode (i.e., highlighting or extraction)
final HTMLHighlighter hh = HTMLHighlighter.newExtractingInstance();
PrintWriter out = new PrintWriter("/tmp/highlighted.html", "UTF-8");
out.println("<base href=\"" + url + "\" >");
out.println("<meta http-equiv=\"Content-Type\" content=\"text-html; charset=utf-8\" />");
out.println(hh.process(url, extractor));
out.close();
Cheers
Original comment by [email protected]
on 24 Nov 2011 at 11:01
from boilerpipe.
That's the correct solution (= HTMLHighlighterDemo.java).
Original comment by ckkohl79
on 24 Nov 2011 at 5:44
- Changed state: Done
- Added labels: Type-Other
- Removed labels: Type-Defect
from boilerpipe.
Can anyone tell me how to output JSON
Original comment by waelmiladi
on 26 Sep 2012 at 9:05
from boilerpipe.
Related Issues (20)
- BoilerplateBlockFilter ignores labelToKeep
- [deleted issue]
- Program does not terminate for badly formatted/syntactically incorrect HTML input
- How to use boilerpipe to get some text with a hyperlink from the web page? HOT 1
- Incomplete extraction of text with special characters
- Server returned HTTP response code: 403 for URL (SOLVED) please use this codeline. HOT 2
- Limit the parsing depth of the html parsing to avoid out of memory situations HOT 1
- Extract article from non-english text HOT 1
- Missing Maven 1.2.0
- Xerces for andorid jar file needed HOT 2
- its not working for a news site HOT 1
- Incomplete extraction of article
- Fail to extract main content on some page, get footnote instead
- IllegalArgumentException for many web pages
- Missing ImageExtractor in downloabale 1.2 jar file
- Performance issues with UnicodeTokenizer
- Boilerpipe is conflicting with CyberNeko library HOT 1
- Unsupported content type: null HOT 1
- Different result when using Web Api and the source api?
- How to debug the result?
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from boilerpipe.