Algorithm
Our aim is to provide concise and coherent summaries. While we started with a relatively simple extraction algorithm for general news, we have now lau`nched a bold and ambitious effort to take summarisation technology far beyond its present state. Our in-house research makes innovative use of machine learning and natural language processing, using evaluation metrics that show our technology to be highly effective.
Summly has joined forces with the best people and tech organizations in this field to push the envelope through massive innovation. We think this is just the beginning of what can be done with summarization technology.
Our aim is to provide concise and coherent summaries. While we started with a relatively simple extraction algorithm for general news, we have now lau`nched a bold and ambitious effort to take summarisation technology far beyond its present state. Our in-house research makes innovative use of machine learning and natural language processing, using evaluation metrics that show our technology to be highly effective.
Summly has joined forces with the best people and tech organizations in this field to push the envelope through massive innovation. We think this is just the beginning of what can be done with summarization technology.
머신러닝하고, 자연언어 처리, 평가 척도??(먼지모르겠다) 등을 사용해서 만든거같은데
나는 이제야 저런 개념을 '들어'보고 툴밖에 '사용'을 못하는데 (내가 인터페이스 프로그래머다 ㅜㅜ)
대단하네 ㅜㅜ
http://summly.com/
swiftkey 라고 키보드앱도 natural language processing 이랑 machine learning 씀... python library많다고하던데... 한번 찾아보삼
헉
결국 15살 애는 그냥 단어 뽑아서 정리하는 알고리즘이였는데 투자받고 다듬은거네..
이게 가능한건지도 상상이안감, 사람이 글을 읽어도 여기가 핵심문장이다 이런건 좀 갈릴 수 있는데 흠..
News모아줘서 보여주는건 그렇게 어려운일은 아닐텐데 서머리를 해준다는건 좀 힘들꺼같음 (주로 주장문에 쓰이는 단어에 가중치를 줬나..?)
Sorry, this isn\'t rocket science at all.
Standard clustering algorithms (found in any off-the-shelf natural text processing library) and text summation with libots should suffice for most of the heavy lifting.
http://tldr.it/
http://libots.sourceforge.net/
Further, most news articles\' first paragraph is a practical (although you may have not noticed) summary. Coming from NLP, unless you can influence the source and the source being Web, the story should be an 80%-20% in the best case -- and you\'ll work VERY hard to correct the remaining 20%, and you WILL remain with a percentage of content you just can\'t summarize properly.
What would make a difference is a real people-driven summation, not machines (see what voicebunny did for text-to-speech, for example). And yes, it would have been fun to combine the two as well.
아까 어떤횽이 말하대로 어떤사람은 real people-driven을 좋아하네
이미 존재 하는 알고리듬을 수정한듯 ->
http://libots.sourceforge.net/
or a majority of news articles, the first three sentences is usually more than perfect as a summarization method -- feature stories on the other hand are insane (and that\'s why we have things like Lexical Chains:
http://acl.ldc.upenn.edu/W/W97/W97-0703.pdf)
미국 게시판 의견들은 대략 1. 아직 존나 구리다. 2. high valuation(높은 가격)은 리서치랑 일하는 엔지니어들때문이다 그리고 제일 많은건 3. 이미 부모백이 존나 쎄고 투자자들 빽이 썠다
오늘도 실망하고 갑니다 요약감사요